2018-10-11

#DSESummit2018, day 2

It is such a great meeting, this meeting. And I think it is because we spent a lot of time early on in this project in building community. That is, we made sure we feel like we are part of a greater whole. Learning from this, I would love to try to bring this community-first thinking to all the things I do. It requires attention!

In the middle of the day, the core team on the project met with the funding officers and we discussed the ramp-down and close-out of the grant. This has two important and very difficult aspects. The first is that we need to finish what we started: The project is to learn about how to do interdisciplinary things in the university, and to communicate successes and failures to other universities and the larger world. I have a role in that and I agreed to take on some of this final communication. The second is to take the best things we are doing in this funded project and figure out how to continue them after the funding is no longer flowing from these granting agencies. That's critical to our success at the NYU CDS. I left the meeting energized, but a bit concerned about what I need to do in the next year or so!

After tremendously interesting discussions and talks, the day ended with a brainstorming session with Richard Galvez (NYU) about possible projects that bring machine learning to the Gaia data. We worked through some simple ideas that I have been thinking about. I like the idea of modeling the Gaia data with deep learning, because even a deep network acting on such small (per-star) data will be tractable, and maybe even interpretable! We ended on optimism, but not with a final decision about what we are going to do.

2018-10-10

#DSESummit2018, day 1

Today was the start of the annual Moore-Sloan Data Science Environments summit. I led an ice-breaker in which we split into small groups and discussed figures and data visualizations. It's a great community, so it was fun to get started. But as for research: I read and commented on text for Bedell (Flatiron) on the plane, and I worked with Richard Galvez (NYU) on designing a small project that brings machine learning to the Gaia data.

2018-10-09

finishing papers; galaxy morphology regressions

The morning started with a conversation between Eilers (MPIA) and I in which we decided that we will finish our connected papers (first draft anyway) by Friday. I think she will make it! But will I make it? I am going to be strong. We also went through some ideas about testing the assumptions that underly our Jeans model for the Milky Way disk, and what to write about the outcomes of those tests.

Mid-day I had good conversations with Storey-Fisher (NYU) about building pseudo-simulations that make point sets with low-amplitude non-trivial power spectra. We spent an unfortunate amount of time figuring out how the numpy fft module organizes and stores fourier transform data. It isn't trivial!

In the afternoon, Elisa Chisari (Oxford) gave a nice (and pleasantly technical) talk about weak lensing, which evolved into a longer discussion about how we might get more information out of galaxy imaging surveys. I pitched my ideas of thinking about how we might train regression models that can predict dark-matter structure from galaxy morphologies or even better large-scale-structure morphologies. And Chisari has (indirect) evidence that such approaches might be very powerful, because (with simulations) she showed (in the context of intrinsic-alignment contamination of weak-lensing data) that even simple measures of galaxy morphology are expected to be very sensitive to the local gravitational tidal field.

One thing that came up in this discussion is my suspicion that ellipticity is a very blunt tool. I have counter-examples that show that ellipticity is not necessarily the galaxy property most sensitive to the weak-lensing field (in an information-theoretic sense). But we formulated a challenge: Make an adversarial morphology distribution for galaxies such that none of the weak-lensing information in the data is in the galaxy ellipticities. That would be hilarious (or instructive, or both).

2018-10-05

so many things!

Ahhh research. After a rocky morning, it was a great research day. Bedell (Flatiron) may have fully debugged all the bugs we introduced earlier this week when we audited and changed the handling of bad and low signal-to-noise data in the HARPS spectra. Price-Whelan (Princeton), Bedell, and I tentatively planned to run The Joker on all of the public exoplanet-relevant extreme-precision radial-velocity data there is. At a meeting, Tomer Yavetz (Columbia) showed the parts of phase space that are at the boundaries between resonant and regular orbits, and he finds that these regions (if there are disrupting objects on these orbits) produce stellar streams that are not thin but fan out chaotically. That delivers some more detailed theoretical understanding of results that Sarah Pearson (Flatiron) obtained and understood a few years ago. Pearson herself is looking at the orbits of the red-giant stars from Eilers (MPIA) and me to see if she can just see the bar, kinematically. Birky (UCSD) and I discussed validation of her results with The Cannon on M-dwarf spectra in APOGEE. She finds that some isochrone models are very consistent with our results, and that we can also estimate stellar radii (which is super-relevant for TESS). Kate Storey-Fisher (NYU) and I broke down what we need to do for our correlation-function estimator to a small set of well-defined sub-projects. Next up: Cheaply simulating weak, Gaussian clustering.

2018-10-04

gravitational wave inferences

Thursdays are low-research days! But I did have a great conversation with Bonaca (Harvard) about the paper we are writing on the GD-1 stellar stream. We talked about the discussion section: What can we say about black-hole models for the gravitational perturbation we observe? What can we say about the population of perturbers from this one perturbing event?

At the end of the day, Will Farr (Flatiron) gave the Departmental Colloquium about gravitational-wave events, with a focus on statistical inference issues. He made some nice points, including that if Advanced LIGO works according to plans, it will generate enough black-hole and neutron-star inspiral events to solve a bunch of cosmological questions, like the Hubble Constant, whether there are pair-instability supernovae and at what masses, and how black-hole binaries form. That is, it will be routine, high-throughput astronomy! Farr is one of the people responsible for the excellent statistical inference underlying the LIGO results.

2018-10-03

more pair-coding; dotastronomy

I got another good pair-coding session in today with Bedell (Flatiron). We had resolved to work on continuum normalization of the HARPS spectra, but instead we ended up working on how to zero-out or delete or censor bad orders and bad epochs of the multi-epoch, multi-order spectra. We came up with simple methods that are hacky but simple and sensible. The whole code seems to be working!

At Stars Meeting, Rocio Kiman (CUNY) told us about her experiences at dotastronomy X, the tenth incarnation of the influential meeting that is the probable origin of hack days, hack weeks, and unconferencing in astrophysics. The short summary is that she loved the meeting and it's culture. Congratulations to the dotastronomy crew, who have changed the world, and Rob Simpson, who started it lo so many years ago.

2018-10-02

power-spectrum estimators

Tuesdays are low-research days, but Kate Storey-Fisher (NYU) and I got to reading the classic FKP paper about how to estimate a power spectrum in a galaxy survey. We think we can do better; maybe much better! But we don't yet understand. Late in the day I mentioned all this to Roman Scoccimarro (NYU) and he gave me some better methods than FKP. I am still optimistic that we have something very very new to say!

2018-10-01

extreme precision radial-velocity; GD-1; TESS

The highlight of my day was a pair-coding session with Bedell (Flatiron) in which we worked through issues with our code wobble that measures radial velocities in extremely high-resolution multi-epoch spectroscopy. The model includes star and telluric models, and regularizations that constrain unconstrained freedoms. The issues are all related to these regularizations: How to set their values, and why various optimization strategies aren't working. We found a few bugs, made a lot of plots, and experimented. In the end: It looks like it is all working! I am so stoked. This could end up being the key project of the Astronomical Data Group at Flatiron. This working session also strongly endorsed (for me, once again) the value of pair coding.

At lunch time I gave the CCPP Brown-Bag talk about the GD-1 projects I am doing with Bonaca (Harvard) and others. It was fun. Several questions from the audience were about what we can understand about the population of perturbers, from this one perturber. That's a good question, to which I have no (current) answer)

Late in the day, I talked to Ben Pope (NYU) about projects in astronomical time-series imaging. He has nice results that show that independent components analysis might be very valuable; this is something that my former student Dun Wang was interested in. And we also discussed things that relate to speckle imaging, lucky imaging, and interferometry. Can we reconstruct good images from many bad ones? And should we? We resolved to do some experiments with the simulated TESS data.

2018-09-28

AstroFest, day 3

Today was the third and final friday of the Gotham AstroFest series, in which we have a very large fraction of the entire astrophysics community in New York City give short talks. This was at NYU, and had contributions from NYU, AMNH, and CUNY scientists. There were a huge number of interesting results in the day. One of the most remarkable things about the day is that fully one quarter of the talks were about black holes. Between NYU and CUNY, there is a lot of research going on related to black holes: Their formation, primordial black holes, their binary dynamics, gravitational-wave signatures, and so on. That's excellent.

A few random highlights for me included: Evidence for weather on brown dwarfs as a function of temperature and gravity by Vos (AMNH), and (relatedly) comparisons between planet and brown-dwarf spectra by Popinchalk (CUNY). It really does appear that there are no strong differences between brown dwarfs and planets (something I discussed with Oppenheimer, AMNH, at lunch). Gandhi (NYU) showed some chemistry and orbits work she has done with Ness (Flatiron) before coming to NYU; that's very related to my interests! Williamson (NYU) visualized a linear SVM, which is beautiful (and old-school). MacFadyen (NYU) convinced us beautifully that his models of the NS—NS merger are really the best!

There was lots on dark-matter detection and dark-matter candidates, including even baryonic and black-hole types. And Tinker (NYU) showed beautiful satellite-galaxy statistics that he got by stacking and background-subtracting galaxy counts in the Legacy Survey imaging for DESI.

If you want to see the full slide deck for the event, it is here.

2018-09-27

how to write a discussion section

In a low-research day, a highlight was a long conversation with Bonaca (Harvard) about the writing of her paper on the GD-1 stream interaction. We discussed structure, and especially the discussion. In a discussion, I like a humble sandwich on proud bread: Start by saying what's most impressive about what we've done, then go into caveats, limitations, approximation wrongness, and the consequences of all that. And then end on a positive note about what kinds of great new things this work will enable going forward.

Late in the day, Alex Kusenko (UCLA, IPMU) spoke about a very wide range of subjects. He claims to have a full explanation for why we don't see the cutoff in the gamma-ray occurrence rate required by photon–photon interactions with the infrared background. He claims that the gamma rays we see from blazars are really reprocessed from cosmic rays. Plausible! But I would need to know a lot more. He also claims to have a way to naturally make primordial black holes in the end stages of inflation, and make all of the dark matter that way. That's interesting. Unfortunately it was such a long and tiring day I couldn't get it together to really check either of these ideas carefully.

2018-09-26

data science for stars; phase space

Our weekly Stars meeting at Flatiron was a pleasure today, as it usually is. Angus (Columbia) and Contardo (Flatiron) are looking at the possibility that we might be able to deblend binary and overlapping stars in the TESS data by their light curves alone. That's crazy, but just crazy enough that I love it! We discussed different ways they might get a training set for this. Luger (Flatiron) asked whether it might be possible to figure out the ell and em (spherical-harmonic order) of the asteroseismic modes by using projections onto transits. That also led to some good discussions about possible methods; many of the crowd liked the ideas that look like lock-in amplification. Marchetti (Leiden) gave us a nice discussion of the high-velocity star results from Gaia DR2. It's too early: The really exciting results will come in data releases 3 and 4 when the magnitude limit for the RVS data gets fainter.

Matt Buckley (Rutgers) showed Adrian Price-Whelan (Princeton) and me his results on measuring phase-space volumes of bound and disrupted objects. The idea is that you might be able to reconstruct the mass of a disrupted object, and say whether it was dark-matter dominated. And get all the attendant dark-matter-theory consequences of that. He showed (unsurprisingly) that observational noise increases the phase-space volume that you naively measure. So we discussed how to approach this. If we are frequentists, maybe we can just ”greedily“ correct the measurements in the direction that lowers the phase-space volume? If we are Bayesians, we have to make more assumptions, I think!

2018-09-25

structure of all models, ever; correlation-function representation

Early in the day I had a long conversation with Leistedt (NYU) about the philosophy of our machine-learning projects. We refined further our view that the machine learning should be part of a larger causal structure that makes sense. My position is that you can think of most (hard) physics problems as having some kind of generalized graphical model with a three high-level boxes. One is called “things I know well but don't care about”, which is things like noise model, instrument model, and calibration parameters. Another is called “things I don't know and don't care about” which is things like foregrounds, backgrounds, and other nuisances. And the last is called “things I don't know and deeply care about”. This last one is our rigid physics model. And the middle one is where the machine learning goes! If we could build models like this very generally, we would be infinitely powerful.

At mid-day, Storey-Fisher and I talked about all the things we could do if we had a correlation function that is not values-in-bins, but was a linear combination of functions. We could look for cosmological gradients. We could do clustering multipoles at small scales, we could estimate the correlation function and power spectrum simultaneously, we could extract Fisher-optimal summary statistics for cosmological parameter estimation. And all these things are possible with our new correlation-function estimator. Next step: Getting the code fast enough to do non-trivial tests.

In the astro seminar at NYU, Savvas Koushiappas (Brown) showed us weak but very interesting evidence that maybe there is a dark-matter annihilation signature in the NASA Fermi data on the Reticulum II dwarf galaxy. Obviously this is incredibly important if it holds up as more data and better calibrations come.

2018-09-24

writing; not ready for TESS

I got some actual writing time in today! I worked on places in the Birky (UCSD) paper (on M-dwarf spectral models) where Birky had left me notes marked "HOGG". That's a great tool: She leaves "HOGG" notes; I search for them in my text editor, and I make the relevant changes or add the relevant text.

Late in the day I had a great conversation with Ben Pope (NYU) about things we can do right now or very soon with TESS artificial data or the first data release of full-frame images. We talked about dimensionality reduction methods, like the robust
PCA methods from Candès and related methods that use convex optimization. We also talked about independent components analysis. In general, when the first data arrive, there will be lots of low-hanging fruit. We also discussed what could be done in advance, with the available artificial data.

2018-09-23

finishing a paper; latents

I dusted off the draft of my paper with Eilers (MPIA) and Rix (MPIA) about spectrophotometric measurements of red-giant distances or parallaxes using Gaia SDSS APOGEE, 2MASS, and WISE. It is nearly done! But we put it on ice while Eilers finished other things. I worked through more than half of the text, making notes on what small things remain to do.

The biggest to-do item? We have a linear model (for the log distance or log parallax or absolute magnitude). That's sweet, because it is simple, and it is interpretable, at least partially. Now we have to make that true by interpreting. Interpreting a linear model is harder than fitting a linear model!

I also had conversations with Storey-Fisher (NYU) about models for the correlation function and Price-Whelan (Princeton) about Milky Way non-equilibrium dynamical models. On the former, we discussed the difference between the correlation function and any particular estimate of the correlation function. It's a bit complex, because I'm not sure there is even agreement in the community about what would be considered the true latent correlation function in the low-ish redshift Universe.

2018-09-21

stream-as-torus; TESS FFIs

I met up early with Price-Whelan (Princeton) to work on the chemical-tangents method papers. This work devolved into rearranging and organizing into categories the to-do list, using GitHub's project tools. That was useful! But it felt a bit like we didn't get anything done. I know that isn't true!

A bit later in the morning we called Jo Bovy (Toronto) to get some advice for Lauren Anderson (Flatiron) on fitting streams in the Milky Way halo. I had been summarizing one of Bovy's papers as saying that streams are close to orbits (that is, you can fit a stream as an orbit) but Bovy corrected us: His paper shows that streams are close to tori. That is, you can expect all the stars in the stream to have similar actions or invariants, but they will not line up as a line on the torus the same way that a single segment of a single orbit would. Duh! That makes good sense and suggests a beautifully simple method for modeling streams with tilted torus sections. I think I almost know how we might do that.

I also checked in with the group working on NASA TESS full-frame images (FFIs), led by Ben Montet (Chicago), who have been hacking at Flatiron all week. They intend to reformat the full-frame images into manageable (and more useful) data objects, extract aperture photometry flexibly, and perform best-in-class de-trending using other stars or other pixels, in the spirit of many things we have done over the years with Kepler data. They really look like a team that might take over the world! For context: The TESS Mission plans to release the raw FFIs with no proprietary period, and they plan to leave it to the community to build open-source (or not!) data-analysis tools around them. Go team!