2018-10-19

target selection; rock and metal

At Flatiron we have purchased a share in the Terra Hunting Experiment, which will be a big, long-term radial-velocity monitoring program with HARPS3. Today Megan Bedell (Flatiron) and I had a conversation about target selection for that survey. There are many choices that could be made in target selection that could make populations or astrophysics inferences very difficult or even impossible later. These conversations remind me of the great and hard work that went in to target selection in the SDSS family of surveys.

The day ended with a great talk by Leslie Rogers (Chicago) about the things that set planet sizes (as a function of mass). She always phrases her results in terms of what isn't rocky, because of the one-sided-ness of some or most of the composition-related observational uncertainties, but it sure looks to my eyes like the smallest planets are rock and metal, like the Earth. She has one extremely good case, which is orbiting so close to its host star that tidal-disruption arguments come in to play! She also was optimistic that transit-timing information might be informative in the near future. There were jokes about water planets and soda-water planets, because many planets that are rich in water are also expected to be very rich in CO2.

2018-10-18

convexity in machine learning

Thursdays are low-research! But there was a great NYU Physics Colloquium at the end of the day by Eric Vanden-Eijnden (NYU) about the mathematical properties of neural networks. I would say “deep learning” but in fact the networks that are most amenable to mathematical analysis are actually shallow and wide.

I am not sure I fully understood EVE's talk, but if I did, he can show the following: Although the optimization of the network (which is a shallow but wide fully connected logistic network, maybe) is not in any sense convex, and although the model is non-identifiable, with certain (or any?) convex loss function, and with enough data (maybe), the optimum of the loss is convex in the approximation of the model to the function it is trying to emulate.

If anything even close to this is true it is extremely important: Can an optimization be non-convex in the parameter space of a function but convex in the function space? I am sure there are trivial examples, but non-trivially? This might relate to things I have wondered about bi-linear models and related, previously.

2018-10-17

bar, spiral structure, and interactions

Stars meeting at Flatiron was absolutely great today. Discussions by Cunningham (UCSC) who has done HST astrometry, Keck spectroscopy, and kinematic analyses of large samples of Milky Way halo stars. She is full stack! And by Brendan Brewer (Auckland) who is working on information theory (in a Bayesian context) to think about experimental design in realistic contexts. My loyal reader knows how close to my heart that is.

Also at Stars meeting, Pearson (Flatiron) and Laporte (UVic) showed models of the effects of the Sagittarius merger on the Milky Way disk. Because the disk is such a sensitive dynamical “antenna”, it should show evidence of this encounter. In the simulations, it appears that the encounter is capable of raising the bar and spiral structure that is very similar to what is observed. Like very similar. This is incredibly exciting: If this pans out, it opens up use of bars and spirals to find or time or weigh galaxy encounters and interactions. Maybe even with dark-matter substructures! Super exciting.

Before all that, Sinan Deger showed me nice results on galaxy morphologies as a function of environment and location around clusters, and Ari Pakman (Columbia) gave a beautiful math-filled talk about Hamiltonian Monte Carlo. He had a very nice, extremely simple proof and picture for why HMC works.

2018-10-16

Is the Milky Way halo really a thing?

A very low-research day was saved by Suroor Gandhi (NYU) who showed me work she is doing with Melissa Ness (Columbia) on stellar chemistry and kinematics. We discussed the question of whether the Milky Way stellar halo really looks like a distinct kinematic and chemical component (as it should!) or whether it just looks like some kind of continuous extension of the disk (which it should not, but does). Interesting, and how to dig deeper?

2018-10-15

stacking residuals?

In an extremely rare event, I finished a paper! Well, a second draft anyway. The plan is to submit next week. This is my paper with Eilers (MPIA) and Rix (MPIA) on spectrophotometric distances.

Other research today included a conversation with Bedell (Flatiron) about how to look at telluric variability in the wobble residuals. In general the residuals are informative! More thoughts about that happened late in the day with Ben Pope (NYU) who had ideas about stacking the wobble residuals in the planet or companion rest frame to find interesting things for different kinds of companions.

And I had a long conversation with Anderson (Flatiron) about applying variational inference to dust or extinction estimates in the Milky Way. We are making a proposal to David Blei (Columbia) and his group to start a collaboration along these lines.

2018-10-14

machine learning; finishing a paper

I worked a bit of the weekend. I had a great conversation with Francois Lanusse (Berkeley) about the uses and abuses of machine learning in astrophysics. We agreed on most things. He sang the praises of some of the newly available cloud services that do machine learning for you. We discussed some pie-in-sky projects.

Months ago, I promised Christina Eilers (MPIA) that when she finished her paper on her Jeans model of the Milky Way disk, I would finish my paper on spectrophotometric parallax (or distance) estimates. Well, today she finished her paper! So I went into panic mode and by the end of the day I was nearly finished. Nearly. I must get up early and finish tomorrow. If I really do finish it, it will be a rare and special thing: A first-author paper! I only write one of those every two or three years.

2018-10-12

#DSESummit2018, day 3

In one of today's lightning talks, Chris Holdgraf (Berkeley) showed us JupyterHub and related projects, which are methods for distributing data, software, and compute to students (or members of a group) so they can transparently use a non-trivial data-science environment, through any kind of client. It is beautiful stuff, but also very interesting in its origins: It grows out of the undergraduate class Data 8 at Berkeley, which is an innovative project to teach the fundamentals of data science to all Berkeley undergrads, independent of their backgrounds. And much later, over drinks, Holdgraf explained to me lots of chaos-monkey-ish and sensible things they do at Berkeley to make sure that their code is truly and absolutely platform-neutral and vendor-independent. The intellectual content of these projects is truly impressive.

In the afternoon, I got some quality time in with Sarah Stone (UW) on our commitments to produce final products for this project around spaces. We discussed the role of ethnography, architects, and data scientists in figuring out what is and isn't working in our spaces. We also discussed what kinds of products we want to produce.

The last event of the day included a great plenary by Huppenkothen (UW) about the AstroHackWeek and related projects. She emphasized its interdisciplinarity, its values of experimentation, and above all, its commitment to broadening the fields of study and being welcoming to all. It was inspiring and enjoyable. I am extremely proud to have been a part of these projects.

2018-10-11

#DSESummit2018, day 2

It is such a great meeting, this meeting. And I think it is because we spent a lot of time early on in this project in building community. That is, we made sure we feel like we are part of a greater whole. Learning from this, I would love to try to bring this community-first thinking to all the things I do. It requires attention!

In the middle of the day, the core team on the project met with the funding officers and we discussed the ramp-down and close-out of the grant. This has two important and very difficult aspects. The first is that we need to finish what we started: The project is to learn about how to do interdisciplinary things in the university, and to communicate successes and failures to other universities and the larger world. I have a role in that and I agreed to take on some of this final communication. The second is to take the best things we are doing in this funded project and figure out how to continue them after the funding is no longer flowing from these granting agencies. That's critical to our success at the NYU CDS. I left the meeting energized, but a bit concerned about what I need to do in the next year or so!

After tremendously interesting discussions and talks, the day ended with a brainstorming session with Richard Galvez (NYU) about possible projects that bring machine learning to the Gaia data. We worked through some simple ideas that I have been thinking about. I like the idea of modeling the Gaia data with deep learning, because even a deep network acting on such small (per-star) data will be tractable, and maybe even interpretable! We ended on optimism, but not with a final decision about what we are going to do.

2018-10-10

#DSESummit2018, day 1

Today was the start of the annual Moore-Sloan Data Science Environments summit. I led an ice-breaker in which we split into small groups and discussed figures and data visualizations. It's a great community, so it was fun to get started. But as for research: I read and commented on text for Bedell (Flatiron) on the plane, and I worked with Richard Galvez (NYU) on designing a small project that brings machine learning to the Gaia data.

2018-10-09

finishing papers; galaxy morphology regressions

The morning started with a conversation between Eilers (MPIA) and I in which we decided that we will finish our connected papers (first draft anyway) by Friday. I think she will make it! But will I make it? I am going to be strong. We also went through some ideas about testing the assumptions that underly our Jeans model for the Milky Way disk, and what to write about the outcomes of those tests.

Mid-day I had good conversations with Storey-Fisher (NYU) about building pseudo-simulations that make point sets with low-amplitude non-trivial power spectra. We spent an unfortunate amount of time figuring out how the numpy fft module organizes and stores fourier transform data. It isn't trivial!

In the afternoon, Elisa Chisari (Oxford) gave a nice (and pleasantly technical) talk about weak lensing, which evolved into a longer discussion about how we might get more information out of galaxy imaging surveys. I pitched my ideas of thinking about how we might train regression models that can predict dark-matter structure from galaxy morphologies or even better large-scale-structure morphologies. And Chisari has (indirect) evidence that such approaches might be very powerful, because (with simulations) she showed (in the context of intrinsic-alignment contamination of weak-lensing data) that even simple measures of galaxy morphology are expected to be very sensitive to the local gravitational tidal field.

One thing that came up in this discussion is my suspicion that ellipticity is a very blunt tool. I have counter-examples that show that ellipticity is not necessarily the galaxy property most sensitive to the weak-lensing field (in an information-theoretic sense). But we formulated a challenge: Make an adversarial morphology distribution for galaxies such that none of the weak-lensing information in the data is in the galaxy ellipticities. That would be hilarious (or instructive, or both).

2018-10-05

so many things!

Ahhh research. After a rocky morning, it was a great research day. Bedell (Flatiron) may have fully debugged all the bugs we introduced earlier this week when we audited and changed the handling of bad and low signal-to-noise data in the HARPS spectra. Price-Whelan (Princeton), Bedell, and I tentatively planned to run The Joker on all of the public exoplanet-relevant extreme-precision radial-velocity data there is. At a meeting, Tomer Yavetz (Columbia) showed the parts of phase space that are at the boundaries between resonant and regular orbits, and he finds that these regions (if there are disrupting objects on these orbits) produce stellar streams that are not thin but fan out chaotically. That delivers some more detailed theoretical understanding of results that Sarah Pearson (Flatiron) obtained and understood a few years ago. Pearson herself is looking at the orbits of the red-giant stars from Eilers (MPIA) and me to see if she can just see the bar, kinematically. Birky (UCSD) and I discussed validation of her results with The Cannon on M-dwarf spectra in APOGEE. She finds that some isochrone models are very consistent with our results, and that we can also estimate stellar radii (which is super-relevant for TESS). Kate Storey-Fisher (NYU) and I broke down what we need to do for our correlation-function estimator to a small set of well-defined sub-projects. Next up: Cheaply simulating weak, Gaussian clustering.

2018-10-04

gravitational wave inferences

Thursdays are low-research days! But I did have a great conversation with Bonaca (Harvard) about the paper we are writing on the GD-1 stellar stream. We talked about the discussion section: What can we say about black-hole models for the gravitational perturbation we observe? What can we say about the population of perturbers from this one perturbing event?

At the end of the day, Will Farr (Flatiron) gave the Departmental Colloquium about gravitational-wave events, with a focus on statistical inference issues. He made some nice points, including that if Advanced LIGO works according to plans, it will generate enough black-hole and neutron-star inspiral events to solve a bunch of cosmological questions, like the Hubble Constant, whether there are pair-instability supernovae and at what masses, and how black-hole binaries form. That is, it will be routine, high-throughput astronomy! Farr is one of the people responsible for the excellent statistical inference underlying the LIGO results.

2018-10-03

more pair-coding; dotastronomy

I got another good pair-coding session in today with Bedell (Flatiron). We had resolved to work on continuum normalization of the HARPS spectra, but instead we ended up working on how to zero-out or delete or censor bad orders and bad epochs of the multi-epoch, multi-order spectra. We came up with simple methods that are hacky but simple and sensible. The whole code seems to be working!

At Stars Meeting, Rocio Kiman (CUNY) told us about her experiences at dotastronomy X, the tenth incarnation of the influential meeting that is the probable origin of hack days, hack weeks, and unconferencing in astrophysics. The short summary is that she loved the meeting and it's culture. Congratulations to the dotastronomy crew, who have changed the world, and Rob Simpson, who started it lo so many years ago.

2018-10-02

power-spectrum estimators

Tuesdays are low-research days, but Kate Storey-Fisher (NYU) and I got to reading the classic FKP paper about how to estimate a power spectrum in a galaxy survey. We think we can do better; maybe much better! But we don't yet understand. Late in the day I mentioned all this to Roman Scoccimarro (NYU) and he gave me some better methods than FKP. I am still optimistic that we have something very very new to say!

2018-10-01

extreme precision radial-velocity; GD-1; TESS

The highlight of my day was a pair-coding session with Bedell (Flatiron) in which we worked through issues with our code wobble that measures radial velocities in extremely high-resolution multi-epoch spectroscopy. The model includes star and telluric models, and regularizations that constrain unconstrained freedoms. The issues are all related to these regularizations: How to set their values, and why various optimization strategies aren't working. We found a few bugs, made a lot of plots, and experimented. In the end: It looks like it is all working! I am so stoked. This could end up being the key project of the Astronomical Data Group at Flatiron. This working session also strongly endorsed (for me, once again) the value of pair coding.

At lunch time I gave the CCPP Brown-Bag talk about the GD-1 projects I am doing with Bonaca (Harvard) and others. It was fun. Several questions from the audience were about what we can understand about the population of perturbers, from this one perturber. That's a good question, to which I have no (current) answer)

Late in the day, I talked to Ben Pope (NYU) about projects in astronomical time-series imaging. He has nice results that show that independent components analysis might be very valuable; this is something that my former student Dun Wang was interested in. And we also discussed things that relate to speckle imaging, lucky imaging, and interferometry. Can we reconstruct good images from many bad ones? And should we? We resolved to do some experiments with the simulated TESS data.