2015-02-11

Penn State, day 2

I had a great lunch today with the graduate students at Penn State. What a great group of students, many of whom are doing things that are sophisticated along directions of computation, hardware, and inference, all of which I love! After lunch I had many conversations, one highlight of which was Alex Hagen (PSU) and I ranting about "upper limits" (argh), and another of which was Runnoe (PSU) showing me very exciting results on possible black-hole–black-hole binaries. She has a set of systems where the broad line is shifted far from the narrow lines, and appears to have a relative acceleration over time. That's exactly what we were looking for with Decarli and Tsalmantza a couple of years ago!

Late in the day I gave my seminar, which was my new one about data-driven models and The Cannon. The questions were great; PSU is a place where the crowd knows all about hardware, statistics, and data analysis! A theme of the questions and answers was that we should be thinking about spectral representations that respect our beliefs about how spectra are causally generated (by absorption lines of finite line-spread function) rather than the straight wavelength-pixel domain. After the seminar, the conversation over dinner with faculty ranged around the modern research university and it's effective love—hate relationship between faculty and administrators!

2015-02-10

Penn State, day 1

Today was my first of two days at Penn State, hosted by Caryl Gronwall (PSU). I had many interesting conversations, too many to mention in detail. Some highlights were the following:

I had lunch with the HETDEX team, who updated me on the project status, and described some of the data properties. We discussed ways you might extract small signals from the enormous numbers of sky fibers that they will have. One point of philosophy: Imaging provides very good photometric measurements; spectroscopy usually does not. Why the difference? It is primarily because images have lots of blank regions where sky can be estimated and very precisely removed, and the point-spread function can be observed. HETDEX will have spectra in a big grid of pixels, so it will have all these properties but in a spectrograph. It might end up making some of the most precise spectrophotometric measurements ever!

Bastien (PSU) and I talked about her photometric-variability method for estimating stellar gravities. We made a prediction that the variability amplitude (which varies non-linearly with logg) might vary linearly with g. She promised to test that and get back to me.

Brandt (PSU) and I discussed the frightening situation with funding, hardware building, projects, and discovery in astronomy. How do we make sure we bring up the next generation of instrument builders if we can't keep funding hardware teams through the standard channels? I have some worries about the future of the profession: Once grant acceptance rates go below some value, the culture (and tenure success rate) might change dramatically, and the profession might lose important skills and people and ideas. We also talked about data-driven models of quasars!

2015-02-09

new planets in K2

Foreman-Mackey gave the brown-bag talk today. He described the method by which he and Ben Montet (Harvard) and others (including me) have found 20-ish new exoplanets in the K2 data. He is writing up the paper now, and fast (I hope).

The key technology is fitting a model for the systematics simultaneously with fitting the exoplanet transit model, for both search and characterization. This is to be contrasted with "fitting and subtracting" a systematics model prior to search. Fit-and-subtract is very prone to over-fitting; most such systems avoid over-fitting by severely restricting the freedom of the systematics model. If you fit the systematics and exoplanet simultaneously, the systematics will not "over-fit" or reduce the amplitude of the (already weak) exoplanet signals, even if the systematics are given a huge amount of freedom (as we give them). If, furthermore, you marginalize out the systematics (as we do), the method is very conservative with respect to the systematics model and search should be close to optimal (inasmuch as your systematics model is a good model for the data).

The upshot is that they find many exoplanets, multiplying the known yield from K2 by a factor of five, and finding some habitable-zone candidates. Also many eclipsing binaries. Very exciting stuff!

2015-02-06

exoplanets from Cambridge, MA

It was great to have John Johnson (Harvard), Andrew Vandenburg (Harvard), Ben Montet (Harvard), and Ruth Angus (Oxford & Harvard) all at CampHogg group meeting today! Each of these brought results to discuss. Vandenburg blew us away with a two-planet system discovered in the K2 data that is in a 3:2 resonance, but so precisely lined up with our line of sight that the transits are almost perfectly coincident in time! Incredible; so unlikely that we discussed the possibility that one of the bodies is an artificial planet placed there by the alien technologists to send us a signal! Johnson showed us an binary-star gravitational lens that is so close, all components can be spectroscopically monitored to compare lens-based inferences with radial-velocity inferences. Perfect agreement!

After show-and-tell, Montet and Foreman-Mackey discussed the state of their K2 search-and-characterization work, and the scope of the first paper, which they spent the rest of the day working on. One thing they mentioned was a brilliant idea from Tim Morton (Princeton) to look for what are known as "astrophysical" false positives: Apply their exoplanet transit depth measurement method not just to the brightness of the star, but also to the x and y position measurements of the star: If the transit appears (gets a finite "depth") in the position measurements, then it is probably a blend with a background star. Beautiful idea.

Along those same lines, we discussed the relationship between the systematics removal of Dun Wang and Foreman-Mackey (using stars to model stars) and that of Vandenburg (using centroid measurements to model stars) and how they are different and the same. I made my counter-intuitive point that the centroid of a star might be encoded at higher signal-to-noise in its brightness than in its actual, direct centroid measurements! This is related to my Kepler-thermometer project idea.

We spent the rest of the day hacking, on exoplanets, writing, and asteroseismology.

2015-02-05

exoplanets and stars

John Johnson and Ben Montet arrived from Harvard today and Tim Morton arrived from Princeton for serious hacking, and Johnson gave our physics colloquium. He spoke about how exoplanets are found and characterized, and the point that if you want to characterize the exoplanets precisely, you must also characterize the stars precisely. Amen! We recognized at the end of the day that we should be talking about The Cannon but didn't get a chance to connect on that. Tomorrow is hacking day, with his group (Exolab) and my group (CampHogg) all in one room!

2015-02-04

doubly intractable group meeting

At group meeting, Huppenkothen introduced us to methods for sampling "doubly intractable" Bayesian inference problems. The problem (and solution) in question is a variable-rate Poisson problem, where you have Poisson-distributed objects (like photons) arriving according to a mean rate that is varying with time, where that rate function is drawn from another process, in this case a Gaussian Process (taken through a function to make it non-negative). The best methods at the present day involve instantiating a lot of additional latent variables and then doing something like Gibbs sampling in the joint distribution of the parameters you care about and the newly introduced latent variables. We didn't understand everything about these complicated methods, but one of the authors, Iain Murray (Edinburgh) will be visiting the group next month, so we plan to make him talk.

Angus arrived for a few days of hacking and we talked about our super-Nyquist asteroseismology projects. We also started email conversations with the authors of this paper and this paper, both of which are impressive for their pedagogical presentation (as well as their results).

2015-02-03

astrohackny, day 1

Today was the first day of Tuesday-morning hacking (as #astrohackny) up at Columbia, with a large crew, led by Adrian Price-Whelan. We discussed what we want to accomplish, what format to adopt, and what data-sets we might play with. The idea is to get work done, learn things, and hack as a group. Last year we did something similar as #nycastroml. Of the data sets we chose to make our principal, agreed-upon hacking data sets, I am most excited about the Planck data. Unfortunately, I am going to miss the next two meetings, but I have high hopes! We start by introducing the data sets and then we will move to pitching, refining, and executing projects on those data sets. We will also spend time getting our own personal projects done; it is partially a "parallel working" time.

On the train back downtown, I had a great conversation with Huppenkothen, Walsh, Vakili, and Ryan about scientific programming and programming languages. We discussed the conditions under which it would make sense to change languages (for example to Julia, the new language of hipness). I argued that there will never be a time during which it is obvious what language to be working in, especially if your data analysis is in any way cutting edge. I also argued that performance matters, even if you are only going to run your code once: The development cycle is unbearable when code is slow!

2015-02-02

frequencies above Nyquist

As my loyal reader knows, I am getting all interested in measuring stellar oscillations—asteroseismology—in data that have integration times (or sampling intervals) too long. For example, G dwarfs have oscillation periods in the 5-minute range, whereas the Kepler data is (by and large) 30-min exposures on 30-min centers. The Kepler data are typical for astronomy, but perhaps not typical examples for "Nyquist sampling" problems, in part because the exposures are integrations (or projections or finite-time averages) rather than samples of the stellar time series that we care about, and in part because the finite-diameter spacecraft orbit makes the periodic-in-spacecraft-time sampling aperiodic in barycentric time.

The integration point hurts me (it attenuates the amplitude of the super-Nyquist signal) but the aperiodicity helps. I discovered today, however, that I don't even need the aperiodicity: All I need is a good model (causal model, I probably should say) of how the data are generated by the stellar signal. I find that if I properly model the integration time, I can see the short-period signals in the data (or at least in Kepler-like fake data). This isn't surprising; it is like "side-band" frequency information in the standard Nyquist case. The key idea behind all this is that we are not ever going to take a Fourier transform or a Lomb-Scargle periodogram; these tools give you the frequencies in the integration-time-convolved stellar signal. We (Angus, Foreman-Mackey, and I) are going to model the stellar signal prior to convolution with the exposure time window.

2015-01-31

cython and fastness

Yesterday I put some stolen time into looking at whether we can measure short-period stellar oscillations in long-cadence Kepler data. The point is that Kepler long-cadence data has 30-min exposure times, but stellar oscillations in G dwarfs have 5-ish-min periods. The aperiodicity of Kepler exposures might save the day: In principle the exposing is drifted relative to a periodic exposing by the light-travel-time variations to the Kepler field induced by the Kepler spacecraft orbit. I wrote code to forward-model this and see if we can infer the very short stellar-oscillation periods. It looks (from preliminary experiments) that we can! If the noise is friendly, that is.

Today I met up with Foreman-Mackey and he showed me how to convert the slow parts of my code over to cython. That sped up my code by a factor of 300. That's why I pay Foreman-Mackey the big bucks! So now I have a functional piece of code that just might be able to do asteroseismology way below the naive "Nyquist limit". Stoked! If we can get this to work, there are thousands (or even tens of thousands) of stars that might present to us their asteroseismological secrets.

2015-01-30

p(z)

At group meeting, Malz discussed this paper by Sheldon et al, which obtains a redshift distribution and individual-object photometric redshifts from photometric data and a heterogeneous training set of objects with spectroscopic redshifts. The Sheldon paper (which refers to the Cunha method) is an example of a likelihood-free inference, in that it creates a redshift distribution and posteriors for redshift for individual objects without ever giving a likelihood function. This is good, in a way, because it doesn't require parameterizing galaxy spectral energy distributions. But it is bad, because, for one, any user of these probabilistic outputs wants likelihood information not posterior information! And more bad because we do have ideas about the probability of the data—we know things about the noise in photometric space—that are the principal inputs to a likelihood function. Malz and I plan to write some kind of paper about all this, but we are still confused about exactly what that paper would contain.

2015-01-29

AAAC, day 2

Today was the second day of the AAAC meeting in DC; I participated remotely from NYC. The most interesting discussions of the day were around proposal pressure and causes for low proposal acceptance rates (around 15 percent for most proposal categories at NSF and NASA). The number of proposals to the agencies is rising while budgets are flat or worse (for individual-investigator grants at NSF they are declining because of increasing costs of facilities). The question is: What is leading to the increase in proposal submissions? It is substantial; something like a factor of three over the last 25 years. At the same time, the AAS membership has increased, but only by tens of percent.

With data in hand, we can rule out a few easy hypotheses: It is not people just sending in lots of bad proposals; proposals are still very good and now proposals rated Excellent and Very Good have less than 50-percent acceptance rates. It is not proposals by young people; the increase seems to be coming from senior people (people with tenure). It is not more proposals in any one year by the same PI; most PIs put in only one proposal to (say) the NSF AAG call in a year. What it might be is that people are not waiting three years between proposal submissions; I know I don't! 25 years ago, a successful research-active astronomer would put in one proposal every three years. Now I think most of us put in a proposal most years. That might account for it, but we don't yet know. It is not absolutely trivial to get the data.

In addition to getting data from the funding agencies, we also plan to do some kind of survey of the community. This will give less "hard" data, but will be able to give answers to questions that we can't ask of the data at the agencies, about things like motivation, perception, and reasoning among proposal PIs.

2015-01-28

AAAC, day 1

Today was a meeting in Washington, DC of the Astronomy and Astrophysics Advisory Committee that advises NSF, NASA, and the DOE on their overlapping interests in astronomy and astrophysics. A few highlights were the following:

NASA plans to start studies of some of the top mission concepts for big missions that might be put forward to the decadal survey in 2020. That is, they want the community to go into the 2020 decadal process with some really well researched and feasible, ambitious missions. The white-paper describing all this is here (PDF) and deserves reading and comment by the community. The best way to influence one of the mission studies is to join and get involved in the relevant "PAG" (whatever that is).

NSF and NASA did well in the FY15 budget process, but DOE HEP less well; this pressures DOE on its priorities. Nonetheless, it is attempting to move forward in sll three areas of next-generation CMB experiment, next-generation dark-matter detection experiment, and something like DESI. The DOE feels a strong commitment in all these areas.

Dressler is chairing a committee to assess the "lessons learned" from the last decadal process. Somehow this slipped past me un-noticed! I wrote him email during the meeting and will send him comments tomorrow, but in fact I think it is too late to influence the survey of the survey, which is too bad. As my loyal reader (or friend) might know, I have issues with some of the decadal-survey processes. In particular, I was horrified by the non-disclosure agreement that members of the survey were asked to sign. (That's why I didn't participate.) It said, and I quote:

"When discussing the survey process with anyone from outside the survey committee, a panel member, or a consultant appointed to the survey, you should avoid discussing any matter other than the publicly available information, such as shown on the NRC’s public web pages.

"Information discussed should be limited to presentations and discussions held in open session, materials circulated in open session only, and other publicly accessible information. Discussions by the committee, subcommittees, and panels held in closed session are confidential and should not be discussed or revealed. Similarly draft documents circulated in closed sessions of a meeting or in the time periods between meetings should be treated as privileged documents and should not be shared outside the survey. This is particularly important for any documents that contain draft conclusions or recommendations, since these survey outputs are not final until the NRC review process for the report is complete."

Astronomers will either honor this agreement—in which case we will never be able to know what really happened in the decadal survey—or else they will violate it—in which case they are violating the law! Indeed, and of course, I know many violations (since who can't talk about what happens behind closed doors?).

On the way home, I built a sandbox for looking at Nyquist-violating Fourier analysis of Kepler data. The idea is (and Vanderplas is a huge supporter of this) that if you have a great generative model of your data, there really is no true "limit". We shall see if we can do anything useful with it.

2015-01-27

K2 search

I spent most of our snow day doing Center for Data Science administration, but when I got a bit of research time, I worked on Foreman-Mackey's K2 exoplanet search paper. I also had a great phone conversation with Rix about The Cannon, and especially what the next papers should be. We need a refactored code and code release paper (led by Anna Ho, MPIA) and we need to try using The Cannon to diagnose issues with the physical models. We also have to use The Cannon to tie together the stellar parameter "systems" of two (or more) surveys, since this was the whole point of the project, originally.

2015-01-26

what can you learn from galaxy spectra?

In the lunch-time brown-bag spot, Mike Blanton gave a great talk about galaxy spectra from SDSS-IV, especially the MaNGA project. He drew an amazingly realistic galaxy spectrum on the board and talked about what can be learned from various spectral features, including galaxy age (or star-formation history), chemical abundances (including abundance anomalies that connect back to star-formation history), and mass function (or dwarf-to-giant ratio). The MaNGA survey will make all of these things possible as a function of position within the galaxy for ten-thousand galaxies, which is exciting. After the talk we discussed a bit what might happen if SDSS-IV APOGEE-2 took some long-exposure spectra of galaxies; presumably there would be tons of dwarf-to-giant information, and also a great measurement of the velocity distributions of stars of different types. Interesting to think about!

2015-01-25

a little writing

In a work day that ended up being almost all not research, I did do a tiny bit of writing in my MCMC tutorial and some cosmological data-analysis fantasies.