2019-06-14

what can you learn from a stream around another galaxy?

My day started with Ana Bonaca (Harvard) telling me about an external-galaxy stream, a stream around an external galaxy found by the Dragonfly telescope. She was able to make a nice model of it! And that fit required that the dark-matter halo be flattened (which is interesting). But what we discussed was (you guessed it) information theory: What can you learn about an external galaxy from seeing a stream in imaging? And what if you get a few radial velocities along that stream? This is a great set of questions, and builds on work we have been doing since the pioneering papers of Johnston and Helmi so many years ago.

2019-06-13

SDSS-V review, day 2

Today was day two of the SDSS-V Multi-object spectroscopy review. We heard about the spectrographs (APOGEE and BOSS), the full software stack, observatory staffing, and had an extremely good discussion of project management and systems engineering. On this latter point, we discussed the issue that scientists in academic collaborations tend to see the burdens of documenting requirements and interfaces as interfering with their work. Project management sees these things as helping get the work done, and on time and on budget. We discussed some of the ways we might get more of the project—and more of the scientific community—to see the systems-engineering point of view.

The panel spent much of the day working on our report and giving feedback to the team. I am so honored to be a (peripheral) part of this project. It is an incredible set of sub-projects and sub-systems being put together by a dream team of excellent people. And the excellence of the people cuts across all levels of seniority and all backgrounds. My day ended with conversations about how we can word our toughest recommendations so that they will constructively help the project.

One theme of the day is education: We are educators, working on a big project. Part of what we are doing is helping our people to learn, and helping the whole community to learn. And that learning is not just about astronomy. It is about hardware, engineering, documentation, management, and (gasp) project reviewing. That's an interesting lens through which to see all this stuff. I love my job!

2019-06-12

SDSS-V review, day 1

Today was day one of a review of the SDSS-V Multi-object spectroscopy systems. This is not all of SDSS-V but it is a majority part. It includes the Milky Way Mapper and Black-Hole Mapper projects, two spectrographs (APOGEE and BOSS), two observatories (Apache Point and Las Campanas), and a robotic fiber-positioner system. Plus boatloads of software and operations challenges. I agreed to chair the review, so my job is to lead the writing of a report after we hear two days of detailed presentations on project sub-systems.

One of the reasons I love work like this is that I learn so much. And I love engineering. And indeed a lot of the interesting (to me) discussion today was about engineering requirements, documentation, and project design. These are not things we are traditionally taught as part of astronomy, but they are really important to all of the data we get and use. One of the things we discussed is that our telescopes have fixed focal planes and our spectrographs have fixed capacities, so it is important that the science requirements both flow down from important scientific objectives, and flow down to an achievable, schedulable operation, within budget.

There is too much to say in one blog post! But one thing that came up is fundraising: Why would an institution join the SDSS-V project when they know that we are paragons of open science and that, therefore, we will release all of our data and code publicly as we proceed? My answer is influence: The SDSS family of projects has been very good at adapting to the scientific interests of its members and collaborators, and especially weighting those adaptations in proportion to the amount that people are willing to do work. And the project has spare fibers and spare target-of-opportunity capacity! So you get a lot by buying into this project.

Related to this: This project is going to solve a set of problems in how we do massively multiplexed heterogeneous spectroscopic follow-up in a set of mixed time-domain and static target categories. These problems have not been solved previously!

2019-06-11

words on a plane

I spent time today on an airplane, writing in the papers I am working on with Jessica Birky (UCSD) and Megan Bedell (Flatiron). And I read documents in preparation for the review of the SDSS-V Project that I am leading over the next two days in a Denver airport hotel.

2019-06-10

spiral structure

This morning on my weekly call with Eilers (MPIA) we discussed the new scope of a paper about spiral and bar structure in the Milky Way disk. Back at the Gaia Sprint, we thought we had a big result: We thought we would be able to infer the locations of the spiral-arm over-densities from the velocity field. But it turned out that our simple picture was wrong (and in retrospect, it is obvious that it was). But Eilers has made beautiful visualizations of disk simulations by Tobias Buck (AIP), who shows very similar velocity structure and for which we know the truth about the density structure. These visualizations say that there are relationships between the velocity structure and the density structure, but that it evolves. We tried to write a sensible scope for the paper in this new context. There is still good science to do, because the structure we see is novel and spans much of the disk.

2019-06-06

information theory and noise

In my small amount of true research time today, I wrote an abstract for the information-theory (or is it data-analysis?) paper that Bedell and I are writing about extreme-precision radial-velocity spectroscopy. The question is: What is the best precision you can achieve, and what data-analysis methods saturate the bound? The answer depends, of course, on the kinds of noise you have in your data! Oh, and what counts as noise.

2019-06-05

alphas, robots, u-band

In an absolutely excellent Stars and Exoplanets Meeting, Rodrigo Luger (Flatiron) had everyone in the room (and that's more than 30 people) say what they plan to get done this summer!

Following that, Melissa Ness (Columbia) talked about the different alpha elements and alpha enhancement: Are all alpha elements enhanced the same way? Apparently models of type-Ia supernovae say that different alpha elements should form in different parts of the supernova, so it is worth looking to see if there are abundance differences in different alphas. The generic expectation is that there should be a trend with Z. She has some promising results from APOGEE spectra.

Mike Blanton (NYU) talked about how we figure out how to perform a set of multi-epoch, multi-fiber spectroscopic surveys in SDSS-V. He has a product called Robostrategy which tries to figure out whether a set of targets (with various requirements on signal-to-noise and repeat visits and cadence and so on) is possible to observe with the two observatories we have, in a realistic set of exposures. That's a really non-trivial problem! And yet it appears that Blanton may have working code. I'm impressed, because integer programming is hard.

And Shuang Liang (Stony Brook) showed us that it is possible to calibrate u-band observations using the main-sequence turn-off, as long as you account for the differences between the disk and the halo. He has developed empirical approaches, and he has good evidence that his calibration based on the MSTO is better than other more traditional methods!

2019-06-04

titles, then abstracts, then figures

I had conversations today with Megan Bedell (Flatiron) and Kate Storey-Fisher (NYU) about titles for their respective papers. I am slowly developing a whole theory of writing papers, which I wish I had thought about more when I was earlier in my career. I made many mistakes! My view is that the most important thing about a paper is the title. Which is not to say that you should choose a cutesy title. But it is to say that you should make sure the person scanning a listing of papers can estimate very accurately what your paper is about.

I then think the next most important thing is the abstract. Write it early, write it often. Don't wait until the paper is done to write the abstract! The abstract sets the scope. If you have too much to put into one abstract, split your paper in two. If you don't have enough, your paper needs more content. And unless you are very confident that there is a better way, obey the general principles (not necessarily the exact form) underlying the A&A structure of context, aims, methods, results.

Then the next most important thing is (usually) the figures and captions. My model reader looks at the title. If it's interesting, the reader looks at the abstract. If that's interesting, they look at the figures. If all that is interesting, maybe they will read the paper. Since we want our papers to be read, and we want to respect the time of our busy colleagues, we should make sure the title, abstract, and figures-plus-captions are well written, accurate, unambiguous, interesting, and useful.

So I spent time today working on titles.

2019-06-03

free energy and life

One amusing conversation today was between Ben Pope (NYU) and myself about whether hot stars are more or less likely to host planets with live. We believe (it's not extremely well established yet) that there are more habitable planets around M-type stars than G-type (and there is probably a relatively smooth function of temperature). So why do we live around a G star? Is it because there is more free energy per photon? I have assumed that this is why. But we realized that we can make this argument quantitative. One question that I have is this: Is this argument anthropic? Or is it just the simple observation that Earth hosts life? I think it is anthropic, because it has something to do with whether our place is special.

2019-06-02

responding to referee

I spent the weekend finally finishing a response to referee on my paper on using spectroscopy and photometry to get precise stellar distances. It was a very constructive, helpful, and positive report, so it really is embarrassing that it took me this long to finish. But it's done and we will resubmit on Monday. Somehow in my dotage, it gets hard to do the things I must do for myself. I am motivated to meet my obligations to others, but it is hard to meet those that help primarily me.

2019-05-30

multi-messenger events and software

Today was a one-day workshop on multi-messenger astrophysics to follow yesterday’s one-day workshop on physics and machine learning. There were interesting talks and discussions all day, and I learned a lot about facilities, operations, and plans. Two little things that stood out for me were the following:

Collin Capano (Hannover) spoke about his work on detecting sub-threshold events in LIGO using coincidences with other facilities, but especially NASA Fermi. He made some nice Bayesian points about how, at fixed signal-to-noise, the maximum possible confidence in such coincidences grows with the specificity and detail (predictive power) of the event models. This has important consequences for things we have been discussing at NYU in our time-domain meeting. But Capano also implicitly made a strong argument that projects cannot simply release catalogs or event streams: By definition the sub-threshold events require the combination of probabilistic information from multiple projects. For example, in his own projects, he had to re-process the Fermi GBM photon stream. Those considerations—about needing access to full likelihood information—has very important implications for all new projects and especially LSST.

Daniela Huppenkothen (UW) put into her talk on software systems some comments about ethics: Is there a role for teaching and training astronomers in the ethical aspects of software or machine learning? She gave the answer “yes”, focusing on the educational point that we are launching our people into diverse roles in science and technology. I spoke to her momentarily after her talk about telescope scheduling and operations: Since telescope time is a public good and zero-sum, we are compelled to use it wisely, and efficiently, and transparently, and (sometimes) even explainably. And with good legacy value. That’s a place where we need to develop ethical systems, at least sometimes. All that said, I don’t think much thought has gone into the ethical aspects of experimental design in astronomy.

2019-05-29

#PhysML19 and likelihoods, dammit

Today was a one-day workshop Physics in Machine Learning. Yes, you read that right! The idea was in part to get people to draw out how physics and physical applications have been changing or influencing machine-learning methods. And it is the first in a pair of one-day workshops today and tomorrow run by Josh Bloom (Berkeley). There were many great talks and I learned a lot. Here are two solipsistically chosen highlights:

Josh Batson (Chan Zuckerberg Biohub) Gave an absolutely great talk. It gave me an epiphany! He started by pointing out that machine-learning methods often fit the mean behavior of the data but not the noise. That's magic, since we haven't said what part is the noise! He then went on to talk about projects noise2noise and noise2self in which the training labels are noisy: In the first case the labels are other noisy instances of the same data, and in the second the labels are the same data! That's a bit crazy. But he showed (and it's obvious when you think about it) that de-noising can work without any external information about what a non-noisy datum would look like. We do that all the time with median filters and the like! But he gave a beautiful mathematical description of the conditions under which this is possible (they are relatively easy to meet; it involves the noise being conditionally independent of the signal, as in causal-inference contexts). The stuff he talked about is potentially very relevant to The Cannon (it probably explains why we are able to de-noise the training labels) and to EPRV (if we think of the stellar variability as a kind of noise).

Francois Lanusse (Berkeley) and Soledad Villar (NYU) gave talks about using deep generative models to perform inferences. In the Lanusse talk, he discussed the problem that we astronomers call deblending of overlapping galaxy images in cosmology survey data (like LSST). Here the generative model is constructed because (despite 150 years of trying) we don't have good models for what galaxies actually look like, or certainly not any likelihood function in image space! He showed a great generative model and some nice results; it is very promising. I was pleased that he strongly made the point that GANs and VAEs don't use proper likelihood functions when they are trained, and those problems might lead to serious problems when you use them for inference. In particular, GANs are very dangerous because of what is called “mode collapse”—the problem that you can generate only part of the data space and still do well under the GAN objective. That could strongly bias deblending, because it would put vanishing probability in real parts of the space. So he deprecated those methods and recommended methods (like normalizing flows) that have proper likelihood formulations. That's an important, subtle, and deep point.

After Lanusse's talk, I came to see him about the point that if LSST implements any deblender of this type (fit a model, deliver galaxies as posterior results from that model), the LSST Catalog will be unusable for precise measurements! The reason is technical: A catalog must deliver likelihood information not posterior information if the catalog is to be used in down-stream analyses. This is related to a million things that have appeared on this blog (okay not a million) and in particular to the work of Alex Malz (NYU): Projects must output likelihood-based measurements, likelihoods, and likelihood functions to be useful. I can't say this strongly enough. And I am infinitely pleased that ESA Gaia has done the right thing.

2019-05-28

snail models; platykurtic galaxies

Suroor Gandhi (NYU) made me in real time some really nice plots today of what a swarm of stars in phase space do over time. Her plots are for the vertical dynamics of the disk. The very exciting thing is that she can reproduce the qualitative properties of The Snail (the phase-space spiral in the local Milky Way disk). Now we have to look at dependence on initial conditions, time of evolution, and potential parameters.

Dustin Lang (Perimeter) and I spoke for a bit about our old project modeling simple galaxy profiles with mixtures of concentric Gaussians. He is trying to build a truly continuous, interpolate-able model for all Sersic indices, and of course (well it wasn't obvious to me before today) at Sersic index of 0.5, the galaxy profile is exactly a Gaussian, so a mixture of Gaussians makes no sense. And then at indices less than 0.5, the distribution becomes platykurtic, so it can't be fit with a concentric mixture of positive Gaussians. What to do there. Lang wants to add in negative Gaussians! I want to say “don't go there”.

2019-05-27

proposal writing

I spent the long weekend, including today, working on my proposal (with Bedell) for the NASA XRP (exoplanets) competition called something like “Extreme-precision radial-velocity in the presence of intrinsic stellar variability”. We propose to do things around modeling the spectra in the joint domain of time and wavelength to improve radial-velocity precision. But we also have an information-theoretic or conceptual part of the proposal. Does NASA support conceptual work? We hope so!

2019-05-24

LIGO housekeeping data, MCMC

Yesterday I gave a talk about data science at Oregon Physics. Today I talked about dark matter—on the chalk board. I talked about various bits of vapor-ware that we are doing with ESA Gaia and streams and The Snail. I also talked about the GD-1 perturbation found by Bonaca and Price-Whelan. That was followed by an excellent and fun lunch with the graduate students, in which they interviewed me about my career and science. Oregon Physics has a great PhD cohort.

In the afternoon, Ben Farr (Oregon) and I hived off to a rural brewery to discuss LIGO systematics and MCMC sampling. I have fantasies about calibrating the LIGO strain data using the enormous numbers of housekeeping channels that are recorded simultaneously with the strain. Farr was encouraging, in that he does not believe that big models of this sort have been seriously done inside the Consortium. That means there might be a role for me or for a collaboration that includes me.

On the MCMC front, we discussed a few different sampling ideas. One is a project by Farr called kombine, which is an ensemble sampler that uses the ensemble to inform an approximation to the posterior, which in turn informs the sampling. Another is vapor-ware by me called bento box which hierarchically splits your problem into a tree of disjoint problems until you get to a set of unimodal problems that are individually trivial. I realized while I was talking I could even use HMC with reflection moves to simplify the problem at the hard boundaries of the boxes in the bento box.

On the drive back to the airport, we found that we agreed on the point that no-one should ever compute a fully marginalized likelihood. That was refreshing; Farr is one of the very few Bayesians I know who get this point. It inspired me to spend my time at the airport tinkering with the paper I want to write on this subject.