2018-01-31

SPHEREx workshop, day 2

I got up at 0530 and looked at the participants and schedule for the SPHEREx workshop. I realized that I had prepared precisely the wrong talk yesterday! So I threw away my slides and made completely new slides. It was rushed. I forgot things. But it was still an improvement. I switched from saying things about scientific goals to saying things about technical improvements or extensions that could make the project more capable in respects that would serve the needs of (among other things) stellar science.

I then headed in to the workshop; I could only make it to the second day. I learned so much today. I can't do it justice. Here are some random facts: A lot could be learned about exoplanets if we could get bolometric fluxes for the stars.
I knew this already, I guess, but the prospects for SPHEREx here are excellent, if the project can deliver absolutely calibrated flux densities. There is a mass–metallicity relationship inside the Solar System! The Solar System contains Trojan satellites/asteroids around Neptune, not just Jupiter! There is no model for the zodiacal light in the Solar System that matches the observations to the level of precision that an infrared survey would need to remove or avoid it. The zodiacal light is consistent with being made up of ground up asteroids and evaporated comets! ALMA has observed many debris disks around nearby stars; some of these are angularly huge. The poster child is Fomalhaut, which has a thin, elliptical ring. It's a crazy thing. I learned these things from a combination of Dan Stevens (OSU), Jennifer Burt (MIT), Carey Lisse (JHU), and Meredith MacGregor (Harvard), but that's just a tiny sampling.

At the end of the day there was discussion of calibration, led by Doug Finkbeiner (CfA) and me. I very much enjoy the technical challenges for SPHEREx and the enthusiasm of the team taking them on.

2018-01-30

slides prep

It was a very low-research day! But on the train to Boston, I prepared slides for a short talk at a meeting at Harvard about the SPHEREx mission concept. I wrote about how this cosmology mission (line intensity mapping and large-scale structure) might revolutionize our knowledge of stars in the Milky Way.

2018-01-29

asteroseismic binaries; distances between transients

Simon J Murphy (Sydney) is in town for two weeks of hacking with Dan Foreman-Mackey (Flatiron). On arrival last week, the two of them implemented something I have been wanting to do for a long time, which is use asteroseismic phase shifts to find binary companions (yes, people have done this for a while now) but without binning the data up or ever explicitly measuring any time delays in bins or at times. This week (having solved that) they are looking at radial-velocity predictions from those discoveries, and testing them with HIRES spectra. They teamed up with Megan Bedell (Flatiron) to use her wobble system to make these measurements. All I did was cheer-lead.

In the afternoon, Alex Malz (NYU) and I discussed what we might do in an upcoming LSST transient classification challenge. I am interested in the following question: Say you have two sparsely and irregularly sampled light-curves of two transient events that are intrinsically similar but maybe at different redshifts, and you want to see that they are similar. How do you construct a relevant, useful, and tractable similarity or distance metric? I have lots of ideas; if we can solve this, we might have something to contribute.

2018-01-26

infall

In a low-research day, Megan Bedell and I went through the assumptions underlying our beliefs about chemical enrichment of a star from infall by dust or rocky material. This is relevant because she is finishing up a paper on the subject, with her extremely precise measurements of Solar twins.

2018-01-25

EPRV, locked planets, data sharing

At lunch-time today, Megan Bedell (Flatiron) and Ray Pierrehumbert (Oxford) gave talks at Flatiron. During Bedell's talk, she nicely laid out the large number of results we have on extreme-precision radial-velocity measurement; we need to start writing papers asap! She even gave a very simple and new description of what we found with respect to HARPS wavelength-calibration fidelity, a couple of years ago. So we need to write that up too.

Pierrehumbert showed fluid-dynamics results on atmospheres of tidally-locked planets (which are interesting, because they sustain huge temperature gradients around their surfaces). He has some cases where he can't find any steady-state solution for the atmophere; the resulting time dependences might have observable consequences.

Late in the day, I gave a presentation to the AAAC that oversees astrophysics and inter-agency cooperation in astronomy across NSF, NASA, and DOE. I was asked to speak about the future of data sharing, data re-use, and joint analyses. I drew inspiration from cosmology and went into two of my standard sets of talking points: The first is that we need to be thinking about likelihood functions, and how to share them: Data sets are combined by their (possibly partially marginalized) likelihood functions. The second is that when data get sophisticated or complex, there is no point in releasing it without also releasing the code that made sense of the data in real scientific projects. That is, code and data releases can't really be seen as separate things. And we might not be able to have a data release without having a code release (with appropriate licensing for repurposing and re-use). My slides were incomplete, but I put them up here anyway.

2018-01-24

get humans out of the loop for EPRV

At breakfast I went off on my complaint that if you have humans integrated into the operational decision-making of any astronomical project, you can destroy the statistical or legacy value of your data. This is a complaint I have had about some extreme-precision radial-velocity projects: Some investigators make decisions, before or during an observing run, to maximize planet yield that involve people in a room, talking. And they are talking about the outcomes of previous observations. That leads to sample selections that depend on the data in complex and un-model-able ways. As in: You would need a complete simulation of the human investigators and their group decision-making even to simulate the observing process, let alone build a probabilistic model of it. That's been a disaster for radial-velocity experiments, and one of the reasons that most of the populations inferences for exoplanets have come from Kepler. Of course, in their defense, many of the radial-velocity projects were trying to maximize planet yield, with statistics be damned!

This made me realize: We should look at things in the experimental-design and active-learning fields to see if we could obtain all the planet-yield benefits that exoplanet projects get from adaptively changing their observing mid-stream, but retain the usefulness of the data for populations inferences, by going to simple algorithmic replacements for the proverbial smoke-filled room. My intuition is a big yes. Another intuition is that any adaptive observing algorithm will necessarily be tuned to a science goal, so you will have to decide in detail what you value in addition to planet yield.

At Stars Group Meeting, John Brewer (Yale) described in detail the EXPRES radial-velocity spectrograph that is being built and commissioned for the Discovery Channel Telescope. He described design decisions, with many valuable contextual comments. One of these is that gas-cell spectrograph designs seem to be declining in popularity, while rigidly controlled lab-bench designs (so calibration parameters vary little and slowly) are rising in popularity. There are many radial-velocity projects getting started right now, so the landscape for this kind of work is changing fast. (Hence the above comments about experimental design!)

2018-01-23

galaxies and dark matter

In another low-research day (it is beginning of term here) there was a great Astrophysics Seminar at NYU by Marilena Loverde (Stony Brook). She has the great capability of turning complex ideas, that are backed by cosmological simulations, into very simple conceptual arguments. And she has been doing this well for a decade now! She showed us that the bias—or the relationship between galaxies and the dark-matter field—cannot be a function of only local density (on some scale). It must also depend on the local expansion history, which depends on perturbations in the dark-matter and dark-energy-ish (quintessence) and radiation fields. This all gives me great hope for my anomalies project with Kate Storey-Fisher (NYU).

2018-01-22

ready to resubmit; new latent-variable model

Combined with effort over the weekend, today I finished my revisions (in response to referee) for the MCMC paper I have written with Dan Foreman-Mackey (Flatiron). The paper is on arXiv but we will update it to the new version if it gets accepted. It is a relief to get it done. And the referee comments were very constructive and valuable. What amazes me is that the AAS Journals are willing to publish such an odd paper. I appreciate it, though: The AAS Journals are great journals.

I got a moment in with Christina Eilers (MPIA) to propose (yet another) latent-variable model for stellar spectra. The idea is that stellar spectra are very simple, so the relationship between the latent variables and the spectra could be linear! The relationship between the spectra and the parameters of interest is definitely non-linear, so we use a Gaussian Process to model the relationship between the latents and the labels. It is a hybrid linear–GP latent-variable model, or HLGPLVM! Oh now that's a great name.

2018-01-19

turbulence

A low-research day was saved by a great talk by Blakesley Burkhart (CfA) about turbulence in astrophysical MHD. She made an unassailable argument that if we don't understand turbulence (and we don't), we don't understand almost any astrophysical phenomenon. That is only slightly an exaggeration! And then she talked about a general framework for understanding turbulence: Simulate the hell out of it, build empirical statistics from those simulations, and use those statistics to measure (latent) physical parameters in observed systems (like interstellar molecular clouds and supernova remnants). She showed baby steps towards this ultimate formalism, but, in my way of thinking, this approach asymptotes to something like ABC or likelihood-free inference, in which we try to put posterior constraints on parameters of interest using statistical surrogates to obviate an explicit likelihood function. Right now, this seems like the only approach for something as complicated as turbulent plasmas.

2018-01-18

inferring stellar parameters from light curves

Today at Flatiron, Melissa Ness (Flatiron) showed Dan Foreman-Mackey (Flatiron) and me that she can infer stellar effective temperature and surface gravity (but not metallicity!) from a Kepler light-curve, with a data-driven model (supervised regression). She feature-izes the light-curve into something for the model by taking the auto-correlation function. This is a clever idea, because it removes phase, or time-translation effects, from the data. And in a cross-validation she can show that she can infer temperatures to something like the precision with which they are measured in the training set! So this is extremely promising, and an interesting extension of things we have seen previously with stellar jitter. I opined that this could only get better if she looks into feature engineering; the autocorrelation function on a uniform grid cannot be the best feature set in any sense. But with encouragement from Foreman-Mackey, we decided to move ahead with a project now and leave feature engineering to later work.

2018-01-17

Gaia and exoplanets

At this morning's Gaia DR2 prep workshop (parallel-working meeting), we gathered a group of people to discuss ideas for using Gaia DR2 data in the service of exoplanet science. We were focusing on easy ideas that could be executed quickly after the data release. These ideas fell into some broad categories. One is to use the astrometry and photometry to get stellar radii and thereby get better estimates of planet radii for the Kepler planets. Another is to use Gaia-based stellar age estimates to compare planetary systems around stars of different ages. Or compare ages for planetary systems of different architectures (as they say). One of my favorite age estimates is the (square of the) vertical action in the Milky-Way disk! Along the same lines: Test theories for pumping or damping of eccentricities with stellar ages. Because of recent work around Flatiron, there was substantial talk of whether Gaia could detect signatures of stars that have recently accreted their planets. There might be different signatures on different time-scales.

Late in the day, I worked some more on my #hackAAS project to look at the dimensionality of stars in element-abundance space. I think (but am not sure) that the best way to think about this problem for the purposes of chemical-tagging applications is in terms of how well we can predict unmeasured abundances. Because this gets at the value or trade-offs between measuring more elements or measuring the elements you already have but better. I need to write first and data-analyze second, but the data are so fun to play with!

2018-01-16

statistics arguments

Today Taisiya Kopytova (ASU) showed up to write up our results on element abundances and binarity in the APOGEE survey. My job is to write up some philosophy of the project. The methodology is extremely clean because we have structured it like a properly executed drug trial. I'm excited to propagate this methodology, that is only really possible in big data sets.

Alex Malz (NYU) and I started our discussion of how to finish his project on combining noisy photometric redshifts to obtain the full redshift distribution. It is a hard paper to write because it has a lot of non-trivial conceptual aspects to it. In particular, there is an argument in the literature that you can just stack the posterior PDFs that is wrong (it's provably correct and yet wrong!), and we have to rebut that without getting too much into the weeds.

2018-01-12

#hackAAS number 6

Today was the 6th annual AAS Hack Day, at #AAS231 in Washington DC. (I know it was 6th because of this post.) It was an absolutely great day, organized by Kelle Cruz (CUNY), Meg Schwamb (Gemini), and Jim Davenport (UW & WWU), and sponsored by Northrup Grumman and LSST. The Hack Day has become an integral part of the AAS winter meetings, and it is now a sustainable activity that is easy to organize and sponsor.

My hack for the day was to work on dimensionality and structure in the element-abundance data for Solar Twins (we need a name for this data set) created by Megan Bedell (Flatiron). I was reminded how good it is to bring data sets to the Hack Day! Several others took the data to play with, and Martin Henze (SDSU) went to town, visualizing (in R) the covariances (empirical correlations) among the elements. Some of these correlations are positive, some are negative, and some are tiny. Indeed, his analysis sort-of chunks the elements up into blocks that are related in their formation! This is very promising for my long-term goals of obtaining empirical nucleosynthetic yields.

What I did for Hack Day was to visualize the data in the natural (element-abundance) coordinates, and then again in PCA coordinates, where many of the variances really vanish, so the data really are low dimensional (dammit; my loyal reader knows that I don't want this to be true). And then I also visualized in random orthonormal coordinates (which was fun); this shows that the low-variance PCA space is indeed a very rare or hard-to-find subspace in the full 33-dimensional element space. I also visualized some rotations in the space, which forced me to do some 33-dimensional geometry, which is a bit challenging in a room of enthusiastic hackers!

But so much happened at the hack day. There was a project (led by aforementioned Bedell) to make interactive web-based plots of the exoplanet systems, to visualize multiplicity, insolation, and stellar properties. There was a project to find the “Kevin Bacon of astronomy” which was obviously flawed, since it didn't identify yours truly. But it did make a huge network graph of astronomers who use ORCID. Get using that, people! Erik Tollerud did more work on his hack-of-hacks to build great tools for collecting and recording hacks, but he also was working on a white paper for NASA about software licensing. I gave him some content and co-signed it. MIT for the win. Foreman-Mackey led a set of hacks in which astronomers learned how to use TensorFlow, which uses NVIDIA GPUs for insane speed-up in linear-algebra operations. Usually people use TensorFlow for machine learning, but it is a full linear-algebra library, with auto-differentiation baked in.

The AAS-WWT people were in the house, and Jonathan Fay, as per usual at Hack Days (what a hacker!), pushed a substantial change to the software, to make it understand and visualize velocity maps. Another group (including Schwamb, mentioned above) visualized sky polygons in WWT, and used a citizen-science K2 discovery as its test case for visualizing a telescope focal-plane footprint. There were nice hacks with APIs, with people learning to use the NASA Astrophysics Data System API and Virtual Observatory APIs, and getting different APIs to talk together. One hack was to visualize Julia Sets using Julia! It took the room a few minutes to get the joke, but the visualization was great, and very few lines of code in the end. And there were at least two sewing hacks.

None of this does justice: It was a packed room, about 1/3 of the participants completely new to Hack Days, and great atmosphere and energy. I love my job!

2018-01-11

nonlinear dimensionality reduction

Today Brice Ménard (JHU) showed me a new dimensionality-reduction method by him and Dalya Baron. He claims it has no free parameters and good performance. But no paper yet!

2018-01-10

streams and dust

At Stars Group meeting, Lauren Anderson (Flatiron) and Denis Erkal (Surrey) both spoke about stellar streams. Anderson spoke about finding them with a new search technique that looks at proper motions for stars found in great-circle segments; this is being prepared for Gaia DR2. Erkal spoke about constraining the Milky-Way potential using only configurational aspects of streams: If a small stream segment locally don't contain the Galactic Center, there must be asphericity in the gravitational potential.

Late in the day, Yi-Kuan Chiang (JHU) showed me absolutely beautiful results cross-correlating various Milky-Way dust maps with high-redshift objects. There ought to be no correlations, at least in the low-extinction regions. But there are correlations, and it is because the dust maps are all contaminated by high-redshift dust in the extragalactic objects themselves (or objects correlated therewith). He can conclude nice things about different dust-map techniques. We discussed (inconclusively) whether his work could be turned around and used to improve Milky-Way map-making.