Today I finished the zeroth draft of a (first-author; gasp!) paper about linear regression with large numbers of parameters. My co-author is Soledad Villar (JHU). The paper shows how—when you are fitting a flexible model like a polynomial or a Fourier series—you can have more parameters than data with no problem, and in fact you often do better in that regime, even in predictive accuracy for held-out data. It also shows that as the number of parameters goes to infinity, your linear regression becomes a Gaussian process if you choose your regularization correctly. It is designed to be like a textbook chapter so we are faced with the question: Where to publish (other than arXiv, which is a given).
2020-12-12
2020-12-11
stellar surface imaging
Today Rachael Roettenbacher (Yale) gave the CCA Colloquium, on Doppler imaging and interferometry and other methods for imaging stellar surfaces. In conversations with Roettenbacher and others (including Lily Zhao at Yale and Megan Bedell and Rodrigo Luger at Flatiron), I'm coming around to the position that extremely precise radial-velocity measurements of stars will require accurate models of rotating, spotty, time-evolving stellar surfaces. At the end of her talk, there were some discussions about this point: Should EPRV surveys be associated with stellar surface imaging campaigns? Probably, if we want to characterize true Earth analogs!
2020-12-09
how clustered is the DM in the Milky Way halo?
Today David Spergel (Flatiron) came by to discuss the following question: How much do we know—empirically—about fluctuations or clustering in the dark matter distribution in the Milky Way halo? Spergel's idea is maybe to use old or metal-poor halo stars: Since stars form in the centers of their DM halos (we think), the clustering or fluctuations in phase space of the old stars should always be larger than the clustering or fluctuations in the dark matter. I bet that's true! And it's easy to test in standard simulations, I think.
2020-12-08
cross-validation, visualized
I have spent a lot of time in my life advocating cross-validation for model selection. That's sad, maybe? But for many reasons, I think it is much better than computing Bayesian evidences or fully-marginalized likelihoods (FMLs on this site!). Today, for the paper Soledad Villar (JHU) and I are writing, I made this figure, which demonstrates leave-one-out cross-validation. Each curve is a different leave-one-out fit, decorated with that fit's prediction for the left-out point. Instructive? I hope so.
2020-12-07
#NeurIPS2020 tutorial
Today Kate Storey-Fisher (NYU) and I led a tutorial at the beginning of the big NeurIPS machine-learning meeting. Our title was Machine learning for Astrophysics and Astrophysics Problems for Machine Learning. We used our forum to advertise astronomy problems to machine-learning practitioners: We astronomers have interesting, hard problems, and our data sets are free! We made public Jupyter notebooks that download data samples to give the crowd a taste of what's available, and how easy it is to get and munge into form.
Our slides are here and our four notebooks (for four different kinds of data) are here, here, here, and here. We got good audience interactions, including especially a group of great astronomers who helped answer questions in the chats!
2020-12-06
working on NeurIPS tutorial
I spent weekend reasearch time elaborating my code notebooks and my slides for the #NeurIPS2020 tutorial that Kate Storey-Fisher (NYU) and I are doing on Monday. I'm a bit stressed with the last-minute prep, but how else would I ever do this??
2020-12-04
group meeting awesome; color light curves from CoRoT
Today at the Astronomical Data Group meeting (led by Dan Foreman-Mackey) we did our quasi-monthly thing of getting a quick update from everyone who shows up. And 18 people showed up! Everyone gave an update; it was great to see the breadth of activity in the group. One contribution that got me excited was Christina Hedges (Ames, but still part of the Group!), who is looking at ESA CoRoT data. The mission was designed to have some tiny bit of color sensitivity, which makes it possible to look at colored light-curve variations and distinguish causal effects. This builds on work by Hedges to look at tiny point-spread-function changes in NASA Kepler and TESS data to get a tiny bit of color information in those white-light missions. Colored light curves are the future.
2020-12-03
Gaia EDR3
Today ESA Gaia EDR3 dropped! It was a fun day; the data are more precise and less noisy! I'm involved in a few different projects with the data. With Hunt (Flatiron) and Price-Whelan (Flatiron) I am looking at the local velocity-space structure in the disk, and seeing if we can classify features by looking at how they vary spatially around the Solar position. With Eilers (MIT) I am going to update our spectrophotometric distance estimates to APOGEE luminous red giants. With Bonaca (Harvard) I will find out if we can improve the kinematics and orbit identification of stellar streams. None of these projects got very far today, but we did make this visualization!
2020-12-02
preparing for a NeurIPS tutorial
Today Kate Storey-Fisher (NYU) and I got together and parallel-worked on our slides for our big #NeurIPS2020 tutorial next week. Somehow it is easier to work on things in parallel!
2020-12-01
constraining transformations to unit determinant
I learned a lot about linear algebra today! I learned that if you exponentiate a matrix (Yes, matrix exponentiation; if this makes you uncomfortable, think about the Taylor series for exponentiation. Do that with a matrix.), the determinant of the resulting matrix is the exponential of the trace of the exponent matrix. So if you need unit-determinant matrices, you can make them by multiplying together exponents of traceless matrices.
Why do I care about all this? Because Jason Hunt (Flatiron), Adrian Price-Whelan (Flatiron), and I realized yesterday that we need to make some of our transformation matrices volume-preserving. This is in our MySpace project for ESA Gaia EDR3 that finds a data-driven, transformation of phase-space to emphasize velocity structure. And I learned that Jax (the simple auto-differentiation tool for numpy and scipy) knows about matrix exponentiation.
I give thanks to Soledad Villar (JHU) for these insights about linear algebra.
2020-11-30
re-doing spectroscopic distances in EDR3 with high-alpha too
Christina Eilers (MIT) and I discussed our ESA Gaia EDR3 projects today. Our top priority is to re-do our machine-learning (linear regression, really) spectrophotometric distance estimates for very luminous red-giant stars, and then re-map the Milky Way disk in abundances and kinematics. We think that even a small improvement in the parallaxes (as we expect to get on Thursday) might make a big difference to our inferred spectroscopic distances. We discussed the point that in our DR2 work we only used stars on the “low-alpha sequence”; we want to generalize if we are going to make complete abundance maps. But also the stars with different abundance trends might want very different distance estimation parameters. That suggests doing the EDR3 regression in a more “abundance-aware” way.
2020-11-27
writing about regression
I spent my research time today (hiding from the world; it's closed here in the US because of Thanksgiving plus pandemic) writing in my paper with Soledad Villar (JHU) about how to perform regression with very flexible models (like polynomials, wavelets, and Fourier modes).
2020-11-25
getting ready for EDR3: cold streams
ESA Gaia EDR3 is next week! I have been trying to get ready in various ways. Today Ana Bonaca (Harvard) showed me remarkable evidence that the cold stellar streams in the Milky Way halo are clustered in kinematic space, maybe also with the halo globular clusters! We discussed things to do in this area. It's perfect for EDR3 because if the clustering is real, it should improve with the improved precision of the EDR3 data.
2020-11-24
how many black holes?
I spoke with Katie Breivik (Flatiron) today about a project to paint toy binary stars (from Breivik's model of how binaries form and evolve) onto toy spectroscopic targets (from Neige Frankel's model of how the Milky Way disk formed) to see how many binary stars and how many black-hole (or compact-object) binaries Adrian Price-Whelan (Flatiron) and I should be finding in the APOGEE survey. The project is simple in principle, but the matching up of differently simulated catalogs is a conceptual and administrative challenge! The hope for this project is that we can constrain something about the formation of black-hole binaries by the observation that we don't find any (or don't find very many) in APOGEE!
2020-11-23
a discriminative model for EPRV
Based on things I have been learning about linear regression this year, I suggested this morning to Lily Zhao (Yale) that she replace our generative model for stellar spectral variability with a discriminative model, which tries to predict the radial-velocity offset of a star from changes in the stellar spectral shape. It's not a trivial model, since the number of parameters (features) is immense (more than 200,000) and the number of training-set examples is small (45-ish). But by the afternoon she did it and it works. Indeed, it looks like it works better than the generative models we have, even when tested by cross-validation.
