2018-04-13

finding planets

In the afternoon today, I spoke to the NYU Physics Majors about how we find planets in the Kepler data. I spoke about the probability calculus, linear models, fitting, marginalization, and search. Nothing here is all that hard; our special abilities are all about the engineering details of doing all that linear algebra fast.

2018-04-12

references

If there is one thing I am worse at than anything else in writing papers, it is properly reading and citing the relevant literature. I spent all morning working through the literature relevant to my Gaia likelihood-function paper. And I know I am still missing things.

2018-04-11

new Hamiltonian sampler; tidal circularization

In our Gaia DR2 prep meeting, I led a discussion of the likelihood function for Gaia data. I made the standard proposal (Gaussian based on Catalog values). And we discussed a bit when you need to use that (that is, be Bayesian) and when can you just treat the data like truth. As my loyal reader might expect, I advised extreme pragmatism.

In Stars meeting, there was lots of great stuff: Foreman-Mackey (Flatiron) showed us his brand-new Hamiltonian MCMC sampler, which has a super-simple interface, and contains lots of the hacky goodness of STAN. He showed that is blows emcee out of the water for some problems with dozens of parameters. He and Price-Whelan (Princeton) are using this sampler to look for multi-star systems in APOGEE radial-velocity data.

Price-Whelan showed us APOGEE data that show very clear evidence of tidal circularization for (fainter) stars orbiting red-giant stars, and in very good agreement with the theory of this. The situation is easiest to predict, apparently, for giant stars with large convective envelopes. He did both the data analysis and some nice theory with MESA and tidal-circularization differential equations too to support his claims.

Late in the day, Julianne Dalcanton (UW) consulted with Foreman-Mackey and me on a hierarchical model for the dust in M33. She wants to simultaneously learn the dust map and the unreddened color-magnitude diagram as a function of position in the galaxy. That's a great but hard optimization.

2018-04-10

undergraduate projects

It was a low-research day! But I did get in a short conversation with Shiloh Pitt (NYU) who is using a Jupyter notebook to verify a set of matrix identities that I am planning on posting to arXiv in the near future. The notebook is great for making outward-facing code! And I also got in a short conversation with Elisabeth Andersson (NYU) who figured out how to make a data-driven model template for searching for new planets in a star that already has known planets.

2018-04-09

Hawking radiation, Gaussians

At lunch time, Matt Kleban (NYU) gave a nice overview of the simplest arguments that black holes must radiate. It was a memorial, of sorts, for Stephen Hawking. Fundamentally, the argument is that if GR is to be consistent with quantum mechanics, then you must have black-hole radiation. That argument is good and sensible, but it is certainly theoretically prejudiced, since it confidently predicts something that will never be observed, and by privileging one part of theory over another. In the discussion afterwards, we learned that GR people tend to think that black holes obviously destroy information, whereas particle physicists tend to think that information will be preserved by some heretofore unknown mechanism. That's interesting, and highlights how socially constructed some aspects of theory might be. But I learned a lot and loved the talk and the discussion. Kleban is a very deep person and a great colleague.

Earlier in the day, I got challenged on a claim that the prior prediction for a snapshot of the amplitude of one mode of a Gaussian-driven damped, harmonic oscillator would be zero-mean and Gaussian. Not the squared amplitude but the straight linear amplitude of the sinusoid with a particular phase. That rattled around in my head all day. Late in the day, I think I have a good argument: Every linear projection of a Gaussian process onto any basis function or anything else (so long as it is a linear function of the Gaussian-process data) will be Gaussian-distributed.

2018-04-07

radial migration

Hans-Walter Rix (MPIA) has been kicking around ideas for observationally testing the process of radial migration of orbits in the Milky Way disk. It is a slippery problem, because stars aren't tagged with their birth locations in the disk! His idea has been to assume that stellar surface metallicity is a nearly deterministic function of time and radius, and then look at explaining all of the variance in abundances we see at any radius today as being the result of radial migration. This is like a maximal approach to the problem. Maximal in the sense that it attributes all variance (in the age-metallicity relationship) to migration.

Today, on the plane home from undislo, I read a very nice draft by Neige Frankel (MPIA) that executes these ideas, and beautifully (and probabilistically). She uses stellar ages from the C and N dredge-up analysis of Melissa Ness (MPIA), which are imprecise but do seem to be ages. Frankel finds sensible parameters for the migration, in terms of radius variance as a function of time. It all hangs together, because the Milky Way data (from APOGEE in this case) really do show a broadening of scatter in metallicity as a function of radius as age increases. That is, there does seem to be at least qualitative support for this picture of radial migration.

2018-04-05

Stitch Fix

I had the great privilege of visiting Stitch Fix today, hosted by Dave Spiegel. I was interested in the company for many reasons, but the main one is that it has a large number of PhD astrophysicists on its data-science team. I learned a huge amount while visiting. Here are some random things:

If you are doing data science to inform or support the decision-making of an employee of your company, it is worth spending a lot of computation on that: After all, the employee is very valuable and expensive! On the other hand, if you are doing data science to directly execute commands (for, say warehousing of goods), you better not get it wrong, because if you have a bug, you could literally move lots of stuff where you don't want it!

If you sell clothing, the lead time between buying clothing wholesale and selling it is long! So you can't quickly or in real time feed back customer preferences into your buying choices. That makes prediction of paramount importance! One thing that really surprised me is that Stitch Fix designs and even manufactures some of its own clothing, so they have unique lines of clothing, adapted to their customers' preferences!

And, obviously (but new to me): Clothing is combinatoric! Even in making a standard button-up shirt, there are myriad few-way decisions about collar, buttons, sleeves, relative dimensions, and so on, such that there is no way in the history of all of humankind that you could make every possible (or even every sensible!) version of a standard dress shirt. That puts a data-science-oriented company like Stitch Fix in a very, very interesting position.

2018-04-03

writing

In an undisclosed location, I continued to write in my Gaia likelihood function document. I'm not sure why! But I'm more convinced than ever that we can deliver likelihood (rather than posterior) information in future versions of (at least some kinds of) probabilistic catalogs.

2018-04-02

a likelihood function for Gaia

I spent time on the long weekend and today working on a short paper or note on the Gaia likelihood function. My point is trivial: It is just to say that we do inference with Gaia by constructing, from the Catalog, a surrogate likelihood function. This is what everyone does but is rarely made fully explicit. I can't tell whether it is worth publishing. But I find myself compelled to write it nonetheless. I guess it makes the additional point that catalogs from projects like Gaia should be based on likelihood information, not posterior information. Why? Because users need to multiply in new likelihoods, not new posteriors, to update their beliefs!

2018-03-29

jackknife, radial migration, and chaos

Yesterday, in conversation with Andrew Mann (Columbia), Jessica Birky (UCSD) and I decided that she should do a full set of jackknife tests on her Cannon model of APOGEE M-dwarf stars. She did that overnight (I love working with such great people!) and the results indeed show that we don't have much good metallicity information about the M-dwarf stars in the training set we have. This inspired Mann to look for more training-set objects; he found a few dozen more, with a bit more metallicity span. Excellent.

In the afternoon, Kathryn Johnston (Columbia) organized a meeting of the Local Local-Group Group. As it were. There were many interesting things discussed. Megan Bedell (Flatiron) showed her Solar-twin abundances and this got a lot of interesting discussion going about their use to constrain Galactic chemical evolution and radial migration in the Milky Way disk. In particular, they could be very constraining if stellar birth composition is a nearly-unique function of time and Galactocentric radius. There were also questions about whether she can constrain nucleosynthetic yields, which I think she can!

Also in that session Tomer Yavetz spoke about chaos and chaotic orbits, and the properties of stellar streams thereon. He had a nice explanation for why chaos shows up so quickly and clearly in stellar streams: The relevant timescale is not the Lyapunov time, but the time it takes for orbits to wander around their local neighborhood in frequency space, which can be a much shorter time (short reason: because that frequency neighborhood can be small). I hope this is correct, because it has been a puzzle!

2018-03-28

ages of M stars

At Stars group meeting, Rocio Kiman (CUNY) showed some beautiful results comparing activity indictors for late-type dwarf stars with kinematic measurements. The stars that are older by activity indicators (less activity) are clearly also older kinematically (higher vertical velocity dispersion). The data are the clearest I have ever seen in this age world. Her goal is to build a hierarchical model of the different age indicators to cross-calibrate them and deliver highest possible precision age measures. She is ready for Gaia DR2!

In Gaia DR2 prep meeting, we went back through our Gaia projects. David Spergel (Flatiron) pointed out to us that the Gaia Archive is fair game for NASA ADAP proposals, which reminded me that I have some serious proposal-writing to do, asap! The discussion of projects got me very excited for April 25, which I think will be a fun celebration of everything astrometric.

2018-03-27

the very local neighborhood

Today Jackie Faherty (AMNH) gave the astro seminar at NYU. She got us fired up about Gaia even before her talk, at lunch, where she said that on April 25 the curtains would finally open and we would get to see the Milky Way for the first time! Her seminar didn't disappoint: She pointed out that of the five closest stars to the Sun, three were discovered in 2014! And it appears that the Solar Neighborhood still has lots of secrets for us to discover. She also showed us a star that passed within 60,000 AU of the Sun some 70,000 years ago. That's interesting! If it disturbed comets onto elliptical orbits, we won't see their infall for a few million years! (Just a free-fall argument there.) That observation, combined with things people have found in Gaia DR1, suggests that we have a close encounter like that about once per million years.

2018-03-26

black holes and quantum neural networks

It seems like a low-research month! But at lunch time, Gia Dvali (NYU) gave us a very surprising black-board talk in which he compared a black-hole horizon (which contains an enormous number of microstates, implied by the black-hole entropy argument) to a quantum neural network with a particular kind of hamiltonian term on each edge. In the network, there is an occupation number for the states in which there is an exponential increase in the number of microstates, which he was arguing is similar to the huge increase in entropy when a black hole forms. That's interesting! But there was plenty of skepticism in the room about its significance. Discussion was heated, especially afterwards.

2018-03-23

model of everything

This morning was the usual parallel-working session at NYU. We discussed various things: Boris Leistedt (NYU) has a draft of a paper that models galaxy photometric data in large-scale structure surveys with a model that includes galaxy types, flexible SEDs for every type, and filter bandpasses, calibration issues or offsets, and luminosity distributions as a function of redshift. That is: A model of everything! Well not quite, but close. This could permit the dream of maximizing information extraction from photometric cosmology surveys, or surveys with mixed photometric and spectroscopic targets. The cool thing is that when the model is causally structured like his, you don't need a representative training set for your photometric redshift estimation.

At the same meeting, Elisabeth Andersson (NYU) showed us a matched filter run over some Kepler data to find a (known) exoplanet, and we discussed how to generalize this to find any other planets that are in a resonant orbit with any of the known planets. Her current plan is to fold and filter, at resonant periods.

Late in the day, Kelle Cruz (CUNY) gave a talk at Flatiron about how to make astronomy better, from an inclusion perspective. Lots of good ideas there; hopefully we can implement them at Flatiron and NYU.

2018-03-22

RV information

My only research today was going through the linear algebra of spectro-perfectionism to see if it is possible that s-p preserves radial-velocity information. I think it doesn't, in the sense that there is more RV information in the two-dimensional spectral image than in the one-dimensional spectrum extracted by s-p. But I don't have a full answer yet. Information theory is hard for me!