I had a great and long conversation today with Maria Bergemann (MPIA) about building a model of a full stellar spectrum out of the models they build of small patches of stellar surface, with full 3D convection and full radiative transfer. Their models are sophisticated, and give a full spectrum in every direction from the surface patch. Thus we can integrate a set of patches into a surrogate combined spectrum for one rotating star covered in convecting patches. We discussed how we might do that, technically, and what projects we might then do with the output. What I want to do (yes, you guessed it) is build data-driven models of stellar granulation to improve radial-velocity surveys.
2023-07-24
2023-07-21
an insight about machine learning
Gaby Contardo (SISSA) completed her visit to Heidelberg today. Over coffee this morning she delivered a very simple, but very nice insight about machine learning outputs. Apologies that this is very Inside Baseball:
As I like to emphasize, you can't really average (or do any populations inferences with) a collection of labels delivered by a discriminative ML method run on a collection of objects. Think: Finding the mean age of a cluster, where each star in the cluster got an age estimate from a discriminative ML method trained on stars with known ages. This is because the discriminative ML methods output something very akin to posterior quantities, and if you average a bunch of posterior estimates, you are multiplying in a prior times itself many times; eventually the prior dominates the inference (in many cases).
Contardo's point: If what you want is a label for a collection of objects, like that mean age, you should train on collections of objects. That is, make a training set where you have sets of N stars, labeled by mean age. Then this model can be applied to a new collection of stars and deliver a mean age estimate! Haha, brilliant. And correct. And consistent with the rules of inference.
2023-07-20
mapping the image plane of a spectrograph
I had a phone conversation about wavelength-calibrating the multi-object APOGEE instrument with Karlo de Leon (NYU) today. He has arc images from each night, and line lists for the arc lamps. But before even using the arc lamps, I recommended that he try to find a model for the 2D images that is an outer product of 1D functions: One is the intensity as a function of wavelength from the arc lamp, and the other is the intensity as a function of slit position from the fibers on the slithead.
The thing we realized in the call is that the coordinate system is right when the image is well described as the outer product of these two functions, warped according to that coordinate system! Okay that's nice, now what will the residuals look like? One issue is that there are cosmic rays, hot pixels, and so on. Another issue is that there will be some vignetting that violates the strict outer-product model. We'll address these issues once we get close.
2023-07-19
extreme infrared excesses
Gaby Contardo (SISSA) showed up in Heidelberg today to make progress on our project on infrared excesses in normal, non-young FGK stars. Because we are using NASA WISE data (along with ESA Gaia and NASA 2MASS), we are only sensitive to bright, hot infrared excesses, much hotter and brighter than typical debris disks around old stars. We have some candidates, which range in temperature from 300 to 1500 K and are reprocessing maybe one percent or a fraction of a percent of the stellar light. (Warning: I haven't calculated this; this is just a guesstimate based on looking at plots.) What are those things? Today we figured out that they can't be warm substellar companions, so they have to be dust (I guess??).
2023-07-18
a likelihood for our Phi-M radio
There are AM radios and FM radios and (if you are a nerd) PCM radios. But Abby Shaum (CUNY) and I have built a Phi-M radio, which demodulates phase variations in a carrier signal. We (with Keaton Bell, CUNY) are using it to find binary companions and planets around stars that show coherent pulsation modes in their photometry. Today I wrote down a noise model for the output of our demodulator. It isn't completely trivial. But it's good, because we can make a likelihood function for fitting our companions. Our model will end up being a limit of the more general model called Maelstrom by Dan Hey (Hawai'i).
2023-07-17
using catalogs responsibly
I had a conversation with Vedant Chandra (Harvard) today about how catalogs are used, and how that relates to how they are built. We started off by arguing about how principled one should be about doing a populations inference. Too abstract! So Chandra moved us in the pragmatic direction: Let's look at a very specific inference and see what matters about it. We decided to look at the distances to distant clusters in the ESA Gaia data: How do your inferences depend on the number of stars you use, the signal-to-noise ratios of those stars, and whether your individual-star measurements are maximum-likelihood or obtained by consideration of a posterior pdf? That should answer questions, and set up some concrete points of discussion.
2023-07-14
data-driven information
My day started with a conversation with Wolfgang Brandner (MPIA), who asked me how to figure out the information content of ESA Gaia RVS spectra, but in a data-driven way. He wants to avoid the theoretical models at first; that is, he wants to figure out how precisely the spectra contain temperature and metallicity and age information without having temperatures, metallicities, and ages that we believe. One approach is to compare to other data that are sensitve to temperature, metallicity, and age: If the RVS spectra can predict those data, then (conditioned on assumptions) they must contain information about temperature, metallicity, and age. This is similar to questions of risk (or expected error in prediction) in machine-learning contexts.
2023-07-13
wobble, star spots, quasar dipole
[Time to try to re-start this forum.]
I spent this morning on three different small activities. One was giving feedback to Matt Daunt (NYU) who is trying to re-build the wobble concept for stellar radial-velocity measurement in jax. He has annoying optimization issues, which are very hard to diagnose! Optimization is always nasty, in my experience.
Another activity was working on the abstract for Lily Zhao's (Flatiron) upcoming paper on stellar variability in the spectral domain, generated by rotating, spotty stars. She is concerned that the paper is too conceptual. I love conceptual papers! I think science moves forward through concepts and implentations, and no individual paper has to do it all.
My third activity this morning was working through the mathematics on a project of Abby Williams (NYU, Caltech) to measure the kinematic dipole in the all-sky Quaia quasar catalog. There are so many different ways to measure it. I think I have a justifiable likelihood function approach, and one in which we could marginalize out—or profile out—the uncertainties in the selection function we have estimated. It's a controversial subject, so I would like to do things correctly.
2023-06-01
co-writing
My only real work time today came at the end of the day, working with Emily Griffith (Colorado) on our paper about element abundance ratios in the SDSS-V APOGEE data. We find (like others) that the abundances can be explained pretty precisely with only two processes. My position (contrary to others) is that this is not because there are two dominant kinds of supernovae! It is because the disk mixes gas quickly in the azimuthal direction, such that a star's properties are mainly set by the (cosmic) time at which it was born, and the radius at which it was born. If I'm right, then you can't really use disk stars to understand process yields, since they are always an ugly mixture of processes. If I am wrong, there's lots more to do!
2023-05-31
Dr Kate Storey-Fisher
Kate Storey-Fisher (NYU) defended her PhD here at NYU today. She killed it! She talked about emulating cosmological simulations (at the level of statistics, not maps), making invariant scalars that encode the shapes and dynamics of dark-matter halos, and her awesome 1.2 million all-sky quasar catalog from ESA Gaia and NASA WISE. It was all things my loyal reader knows lots about but I loved it. It has been an honor and a privilege to work with KSF these years, and I will miss her very very much.
2023-05-30
Dr Irina Espejo
Today it was my honor to serve on the PhD defense committee of Irina Espejo (NYU), who is one of the first (ever in the world, actually!) PhDs in Data Science. Her PhD research involved making real, practical, scalable, reproducible tools for the (late-in-pipeline) analysis of high-energy physics data from the Large Hadron Collider. She built tools to speed up likelihood-free inferences, and she built a tool to find exclusion regions (upper limits) in complex parameter spaces. She used the latter to put constraints on a (real, not toy) proposed modification to the standard model.
On the first project, the tools that she built (and built on) make the LHC more sensitive to new physics, because they find better test statistics for distinguishing models. They make some searches far better, which makes me wonder whether particle physics is using our money efficiently??
2023-05-26
how to extract XP spectra from raw Gaia data?
On the plane home from meetings at Cambridge, Warwick, and Paris, I worked on a long document I am writing for Gaia DPAC CU5, which is the organization responsible for calibrating and extracting the Gaia XP spectra. They are doing a beautiful self-calibration to extract all the spectra on the same system, in the sense of resolution, dispersion, and throughput. But their system has some pathologies, which we discussed last week. I think I know how to solve some of them. My document is reporting those thoughts.
Writing like this reminds me of graduate school: One of my advisors (Blandford) often encouraged me to write up thoughts, ideas, projects, and proposals, even when we had no intention of submitting them anywhere. It's good practice, I think, because you can't understand anything if you don't write about it.
2023-05-25
how to maximize the yield of planets?
There were discussions this week at University of Warwick about the Terra Hunting Experiment strategy and likely detection capability. Various take-homes include that we need to mitigate lots of stellar noise, and that we care deeply about the covariance (as a function of separation in time) of adjacent measurements. I advocated that we split our ten-year survey into two or three surveys, of varying length. In the first, we learn about the stars, and in the last, we go to town on the very most promising targets. There was general agreement that this is a good idea. But now we need a very specific plan for what this means. As my loyal reader knows, in my view, the decisions must be based on repeatable operations, so that we have some hope of learning statistical things about populations in the end.
2023-05-24
predicting RVs from SDO imaging
I'm at the Terra Hunting annual Science Working Group meeting, held this year at University of Warwick. There were many great talks today, some technical and some science. My mind was blown by Ben Lakeland (Exeter), who showed Solar Dynamics Orbiter data of the Sun, and then showed that, from these images, he can predict the magnetic-activity-generated RV signals in simultaneous EPRV measurements of the Solar RV. That's pretty exciting. He also showed that much of the time, the RV variations are dominated not by magnetic activity per se. If we are going to beat one meter per second, we are going to have to correct for convective shifts. Somehow!?
2023-05-23
distances between point clouds
I spent the last two days working at Apple Paris, which was fun! I worked with the open-source ott-jax package, which can do some amazing things. I worked with Soledad Villar (Apple & JHU) to generalize the k-means algorithm to point clouds! It can cluster point clouds morphologically, even if the different point clouds have different numbers of points, and even if the different point clouds live in spaces of different dimensions! Everything obeys permutation and rotation symmetries.