My last research day before heading to MPIA for the summer was taken up with many non-research things! However, I did have brief discussions with Lauren Anderson (Flatiron) about what is next for our collaboration, now that paper 1 is out!
2017-06-29
2017-06-28
spin tagging of stars?
At the Stars group meeting, John Brewer (Yale) and Matteo Cantiello (Flatiron) told us about the Kepler / K2 Science meeting, which happened last week. Brewer was particularly interested in the predictions that Ruth Murray-Clay made for chemical abundance differences between big and small planet hosts; it is too early to tell how well these match on to the results Brewer is finding for chemical differences between stars hosting different kinds of exoplanet architectures.
Other highlights included really cool supernova light curves, with amazing details, Granulation or flicker estimates of delta-nu and nu-max, and a clear bimodality in planetary radii between super-earths and mini-neptunes. There was much discussion in group meeting of this latter result, both what it might mean, and what predictions it might generate.
Highlights for Cantiello included results on the inflation of short-period planets by heating by their host stars. And, intriguingly, a possible asteroseismic measurement of stellar inclinations. That is, you might be able to tell the projection of a star's spin angular momentum vector projected onto the line of sight. If you could (and if some results about aligned spin vectors in star-forming regions hold up) this could lead to a new kind of tagging for stars that are co-eval!
2017-06-27
global ozone
In the morning, researchers from across the Flatiron Institute gathered for a discussion of statistical inference, which is a theme that cuts across the different departments. Justin Alsing (Flatiron) led the discussion, asking for advice on his project to model global ozone over the last few decades. He has data that spans latitude, altitude, and time, and the ozone levels can be affected by many things other than long-term degradation by pollutants. So he wants to build a non-linear, data-driven model of confounders but still come to strong conclusions about the long-term trends. There was discussion of many relevant methods, including large linear models (regularized strongly), independent components analysis, latent variable models, neural networks, and so on. It was a wide-ranging and valuable discussion. The CCB at Flatiron has some valuable mathematics expertise, which could be important to all the Flatiron departments.
2017-06-26
statistics is hard
OMG much of my research time today was spent trying to figure out everything that is wrong with Section 7 (uncertainties in both x and y) of the Hogg, Bovy, and Lang paper on fitting a line. Warning to users: Don't use Section 7 until we update! The problems appeared early (see the GitHub issues on this Chapter), but came to a head when Dan Foreman-Mackey (UW) wrote this blog post. Oddly I disagree with Foreman-Mackey's solution, and I don't have consensus with Jo Bovy (Toronto) yet. It has something to do with how we take the limit to very large variance in our prior. But I must update the paper asap!
2017-06-22
the variance on the covariance of the variance
I had a long set of conversations with Boris Leistedt (NYU) about various matters cosmological. The most exciting idea we discussed comes from thinking about good ideas that Andrew Pontzen (UCL) and I discussed a few weeks ago: If you can cancel some kinds of variance in estimators by performing matched simulations with opposite initial conditions, might there be other families of matched simulations that can be performed to minimize other kinds of estimator variances?
For example, Leistedt wants to make a set of simulations that are good for estimating the covariance of a power-spectrum estimator in a real experiment. How do we make a set of simulations that get this covariance (which is the variance of a power spectrum, which is itself a variance) with minimum variance on that covariance (of that variance)? Right now people just make tons of simulations, with random initial conditions. You simply must be able to do better than pure random here. If we can do this well, we might be able to zero out terms in the variance (of the variance of the variance) and dramatically reduce simulation compute time. Time to hit the books!
2017-06-21
fast bar
Stars group meeting ended up being all about the Milky Way Bar. Jo Bovy (Toronto), many years ago, made a prediction about the velocity distribution as a function of position if the velocity substructure seen locally (in the Solar Neighborhood) is produced (in part) by a bar at the Galactic Center. The very first plate of spectra from APOGEE-South happens to have been taken in a region that critically tests this model. And he finds evidence for the predicted velocity structure! He finds that the best-fit bar is a fast bar (whatever that means—something about the rotation period). This is a cool result, and also a great use of the brand-new APOGEE-S data.
Bovy was followed by Sarah Pearson (Columbia) who showed the effects of a bar on the Pal-5 stream and showed that some aspects of its morphology could be explained by a fast bar. We weren't able to fully check whether both Bovy and Pearson want the exact same bar, but there might be a consistent story emerging.
2017-06-20
MCMC
The research highlight of the day was Marla Geha (Yale) dropping in to Flatiron to chat about MCMC sampling. She is working through the tutorial that Foreman-Mackey (UW) and I are putting together and she is doing the exercises.
I'm impressed! She gave lots of valuable feedback for our first draft.
2017-06-19
learning
I spent time working through the last bits of a paper by Dun Wang (NYU) about image modeling for time-domain astrophysics. I asked him to send it to our co-authors.
The rest of the day was spent in discussions of Bayesian inference with the Flatiron Astronomical Data Group reading group. We are doing elementary exercises in data analysis and yet we are not finding it easy to discuss and understand, especially some of the details and conceptual arguments. In other words: No matter how much experience you have with data analysis, there are always things to learn!
2017-06-16
cosmic rays, alien technology
I helped Justin Alsing (Flatiron) and Maggie Lieu (ESA) search for HST data relevant to their project for training a model to find cosmic rays and asteroids. They started to decide that HST's cosmic-ray identification methods that they are already using might be good enough to just rely upon, which drops their requirements down to asteroids. That's good! But it's hard to make a good training set.
Jia Liu (Columbia) swung by to discuss the possibility of finding things at exo-L1 or exo-L2 (or the other Lagrange points). Some of the Lagrange points are unstable, so anything we find would be clear signs of alien technology. We looked at the relevant literature; we may be fully scooped, but I think there are probably things to do still. One thing we discussed is the observability; it is somehow going to depend on the relative density of the planet and star!
2017-06-15
Bayesian basics; red clump
A research highlight today was the first meeting of our Bayesian Data Analysis, 3ed reading group. It lasted a lot longer than an hour! We ended up going off into a tangent on the Fully Marginalized Likelihood vs cross-validation and Bayesian equivalents. We came up with some possible research projects there! The rest of the meeting was Bayesian basics. We decided on some problems we would do in Chapter 2. I hate to admit that the idea of having a problem set to do makes me nervous!
In the afternoon, Lauren Anderson (Flatiron) and I discussed our project to separate red-clump stars from red-giant-branch stars in the spectral domain. We have two approaches: The first is unsupervised: Can we see two spectral populations where the RC and RGB overlap? The second is supervised: Can we predict relevant asteroseismic parameters ina training set using the spectra?
2017-06-14
cryo-electron-microscopy biases
At the Stars group meeting, I proposed a new approach for asteroseismology, that could work for TESS. My approach depends on the modes being (effectively) coherent, which is only true for short survey durations, where “short” can still mean years. Also, Mike Blanton (NYU) gave us an update on the APOGEE-S spectrograph, being commissioned now at LCO in Chile. Everything is nominal, which bodes very well for SDSS-IV and is great for AS-4. David Weinberg (OSU) showed up and told us about chemical-abundance constraints on a combination of yields and gas-recycling fractions.
In the afternoon I missed Cosmology group meeting, because of an intense discussion about marginalization (in the context of cryo-EM) with Leslie Greengard (Flatiron) and Marina Spivak (Flatiron). In the conversation, Charlie Epstein (Penn) came up with a very simple argument that is highly relevant. Imagine you have many observations of the function f(x), but for each one your x value has had noise applied. If you take as your estimate of the true f(x) the empirical mean of your observations, the bias you get will be (for small scatter in x) proportional to the variance in x times the second derivative of f. That's a useful and intuitive argument for why you have to marginalize.
2017-06-13
Renaissance
I spent the day at Renaissance Technologies, where I gave an academic seminar. Renaissance is a hedge fund that created the wealth of the Simons Foundation among many other Foundations. I have many old friends there; there are many PhD astrophysicists there, including two (Kundić and Metzger) I overlapped with back when I was a graduate student at Caltech. I learned a huge amount while I was there, about how they handle data, how they decide what data to keep and why, how they manage and update strategies, and what kinds of markets they work in. Just like in astrophysics, the most interesting signals are at low signal-to-noise in the data! Appropriately, I spoke about finding exoplanets in the Kepler data. There are many connections between data-driven astrophysics and contemporary finance.
2017-06-12
reading the basics
Today we decided that the newly-christened Astronomical Data Group at Flatiron will start a reading group in methods. Partially because of the words of David Blei (Columbia) a few weeks ago, we decided to start with BDA3, part 1. We will do two chapters a week, and also meet twice a week to discuss them. I haven't done this in a long time, but we realized that it will help our research to do more basic reading.
This week, Maggie Lieu (ESA) is visiting Justin Alsing (Flatiron) to work (in part) on Euclid imaging analysis. We spent some time discussing how we might build a training set for cosmic rays, asteroids, and other time-variable phenomena in imaging, in order to train some kind of model. We discussed the complications of making a ground-truth data set out of existing imaging. Next up: Look at what's in the HST Archive.
2017-06-11
summer plans
I worked for Hans-Walter Rix (MPIA) this weekend: I worked through parts of the After Sloan 4 proposal to the Sloan Foundation, especially the parts about surveying the Milky Way densely with infrared spectra of stars. I also had long conversations with Rix about our research plans for the summer. We have projects to do, and a Gaia Sprint to run!
2017-06-08
music and stars
First thing, I met with Schiminovich (Columbia), Mohammed (Columbia), and Dun Wang (NYU) to discuss our GALEX imaging projects. We decided that it is time for us to produce titles, abstracts, outlines, and lists of figures for our next two papers. We also realized that we need to produce pretty-picture maps of the plane survey data, and compare it to Planck and GLIMPSE and other related projects.
I had a great lunch meeting with Brian McFee (NYU) to catch up on his research (on music!) and ask his advice on various time-domain projects I have in mind. He has new systems to recognize chords in music, and he claims higher performance than previous work. We discussed time-series methods, including auto-encoders and HMMs. As my loyal reader knows, I much prefer methods that deal with the data probabilistically; that is, not methods that always require complete data without missing information, and so on. McFee had various thoughts on how we might adapt methods that expect complete data for tasks that are given incomplete data, like tasks that involve Kepler light curves.