2012-04-09

insane robot model

Imagine you have an insane robot, randomly taking images of the sky. Every so often it takes an image that includes in its field of view a star in which you are interested. When it takes this image, it processes the image with some software you don't get to see, and it returns a magnitude for the star and an uncertainty estimate or else no measurement at all and no uncertainty information at all (because, presumably, it didn't detect the star, but you don't really know). I am in Berkeley this week to answer the following question: Do the null values—the times at which the robot gives you no output at all—provide useful information about the lightcurve of the star?

To make life more interesting, we are assuming that we don't believe the uncertainty estimates reported by the robot, and that the robot returns no information at all about upper limits or detection limits or the data processing, ever. That is, you don't know anything about the non-detections. You can't even assume that there are other stars detected in the images that are useful for saying anything about them.

Of course the answer to this question is yes, as Joey Richards (Berkeley), James Long (Berkeley), and I are demonstrating this week. As long as you don't think that the robot's insane software is varying in sync with properties of your star, the nulls (non-detections) are very informative. We even have examples of badly observed variable stars, for which the use of the nulls is necessary to get a proper period determination.

The crazy thing is, in astronomy, there are many data sets that have this insane-robot property: There are lots of catalogs created by methods that no-one precisely knows, that are missing objects for reasons no-one precisely knows. And yet they are useful. We are hoping to make them much more useful.

2012-04-06

Geha and lucky

Marla Geha (Yale) gave the astrophysics seminar, on the star-formation properties of low-luminosity galaxies. She (with Blanton, Tinker, Yan at NYU) finds an incredibly strong environmental effect: Low-luminosity galaxies that are far from any massive parent galaxy are all star-forming, whereas those more nearby are in a mix of star-forming and non-star-forming. It is not just a trend: There really are no non-star-forming, isolated, low-luminosity galaxies. This makes these galaxies very valuable tools for cosmlogy and galaxy formation, as Geha noted and the crowd discussed. The seminar was just what I love: Lots of constructive and interesting interruptions and discussion, mainly because the results are so interesting and valuable; this really is a significant discovery. It is the strongest (in some sense) known environmental effect for galaxies, and it immediately rules out many simplistic ideas about low-mass galaxies, like that their star formation histories will be shut off by supernova feedback.

On the airplane to visit Bloom's CTDI at Berkeley, I worked more on our side project on Lucky Imaging. I got the code working (finally) and then handed it off to Foreman-Mackey. I have to get back to the things I am supposed to be doing! (Although one of the things that I love about my job is that procrastination like this is valuable in the long term.)

2012-04-05

not much

I didn't get much work done today except some preparation for my week (next week) at the CTDI at Berkeley. I also spend a nice chunk of time discussing induction and inference with Michael Strevens (NYU Philosophy). It turns out that we can make use of each others' expertise.

2012-04-04

lucky sprint

Foreman-Mackey and I (very irresponsibly) did a monster code sprint today on lucky imaging. We got a piece of code working, doing online stochastic gradient like in the Hirsch et al (2011) paper I cited two days ago. We are running on data provided by Federica Bianco (UCSB). However, we have not applied various useful and important regularizations (like centering the PSFs, making the PSFs non-negative, and so on) so currently the code wanders off to fuzzy model space. More soon!

2012-04-03

mixture of Gaussians

I started writing a short note on using mixtures of Gaussians to improve two-dimensional image fitting. I realized in writing it that our expansions of the exponential and de Vaucouleurs profiles in terms of concentric Gaussians have an amusing property: They provide (automatically) their own three-dimensional deprojections, in the optically thin limit. That is, a set of concentric, two-dimensional, isotropic Gaussians that make a deVaucouleurs profile (and we have that set) can be the projection of a very simple-to-compute set of concentric, three-dimensional, isotropic Gaussians. That could be amusing. I should also look at fitting mixtures of Gaussians to various three-dimensional profiles too. This reminds me of my (dormant) project (with Jon Barron) to build a three-dimensional model of all Galaxies. But of course the main point of all this is to make image fitting code fast and numerically stable.

2012-04-02

taking the lucky out of lucky imaging

Federica Bianco (UCSB) stopped by for a few hours; she will join us at the CCPP next year as a fellowship postdoc. We discussed lucky imaging, which she has been doing with the LCOGT project. I bragged to her that I think I can beat any standard lucky imaging pipeline, and now I have to put up or shut up. I was basing my claims on this obscure paper which is packed with awesome, game-changing ideas. The basic idea is that if you see each image as putting a constraint on a PSF-convolved scene, every image, no matter how bad its PSF, contributes information to the high-resolution image you seek. The cool thing is that if you have enough data, the final reconstructed scene will have better angular resolution than even your best single image.

2012-03-30

SVM, ML, hierarchical, and torques

Fadely dropped in on Brewer, Foreman-Mackey, and me for the day. He has very nice ROC curves for the performance of (1) a support vector machine (with the kernel trick), (2) our maximum-likelihood template method, and (3) our hierarchical Bayesian template method, for separating stars from galaxies using photometric data. His curves show various things, not limited to the following: Hierarchical Bayesian inference beats maximum likelihood always (as we expect; it is almost provable). A SVM trained on a random subsample of the data beats everything (as we should expect; Bernhard Schölkopf at Tübingen teaches us this), probably because it is data-driven. A SVM trained much more realistically on the better part (a high signal-to-noise subsample) of the data does much worse than the Hierarchical Bayesian method. This latter point is important: SVMs rock and are awesome, but if your training data are different in S/N to your test data, they can be very bad. I would argue that that is the generic situation in astrophysics (because your well labeled data are your best data).

In the afternoon, David Merritt (RIT) gave a great talk about stellar orbits near the central black hole in the Galactic Center. He showed that relativistic precession has a significant effect on the statistics of the random torques to which the orbits are subject (from stars on other orbits). This has big implications for the scattering of stars onto gravitaitonal-radiation-relevant inspirals. He showed some extremely accurate n-body simulations in this hard regime (short-period orbits, intermediate-period precession, but very long-timescale random-walking and scattering).

2012-03-29

measuring the undetectable

Brewer, Foreman-Mackey, and I have been working hard on sampling all week, so we took a break today to discuss the astrophysics of the problem. We want to measure the number–flux relation (or luminosity function) of stars in a cluster, below where confusion becomes a problem for identifying or photometering stars. There is an old literature from radio astronomy in the sixties and seventies about inferring properties of faint, overlapping sources from looking at statistics of the resulting confusion noise. But nowadays with awesome probabilistic techniques like Brewer's, we might be able to directly generatively model the confused background, and thereby produce probabilistic information about the astrophysical systems that generate it. Could be huge, in these days of confused Herschel data, the PHAT data on M31, the Galactic Center, and so on.

One of the things we discussed is how to make the problem as challenging as possible. We don't want to do any cheating, where the brighter (resolved) stars are effectively giving us the information we seek about the fainter (unresolved) stars.

2012-03-28

sampling and degeneracies

Brendon Brewer (UCSB), Foreman-Mackey, and I worked on sampling for most of the day, with some calls to Lang. I want to start a challenge called MCMC High Society, which is a bake-off for samplers that can handle ten thousand parameters. However, in the benchmark problem we have, Brewer's DNest sampler is crushing emcee! That might not be surprising, because we have chosen a problem with massive combinatoric degeneracies (with large low-probability valleys separating them) and emcee is not awesome in that situation. Insights from the day include the following: You should measure your sampler's autocorrelation time in the data space (prediction space) not the parameter space (at least if you aren't a realist, and I am not). Ensemble samplers (like emcee) that work with pairs of walkers should keep track of the acceptance fraction as a function of the walker pair involved in the proposal. Multi-modality in the posterior is the norm, not the exception, and it is hard to deal with for any sampler; that's the whole deal. Samplers that estimate the Bayes integral as they sample might be slow for easy problems, but they might be much better for hard problems, and it is hard problems that are game changers.

2012-03-27

not much

Can't say I did all that much today, which is a problem since it is research-only-Tuesday. I spoke at length with Lang about Tractor paper 0, which is still just a gleam in our eyes. I spoke briefly with Brewer and Foreman-Mackey, who are working to bake-off samplers on a very hard problem (see yesterday's post). I worked a bit more on a note about photometric redshift ideas for Alexandra Abate (Arizona), but it isn't moving forward very fast!

2012-03-26

Brewer benchmark

Today Brendon Brewer (UCSB) showed up for a week of probabilistic reasoning. He sat in on the weekly meeting of Goodman, Hou, Foreman-Mackey, and I where we discuss all things sampling and Bayes and stellar oscillations (these days). In the afternoon, we tried to get serious about what we would try to accomplish by the end of the week. One thing Brewer and we are both interested in is the performance of samplers in situations where the number of parameters gets large. Another is adapting Foreman-Mackey's emcee code to make it capable of measuring the evidence integral (fully marginalized likelihood).

In the end, we formulated a problem that is amusing to all of us and also a good test of samplers in the large-numbers-of-continuous-parameters regime: Modeling an astronomical image of a crowded field. We figured out three potential publications: The first is a benchmark for samplers working in the astronomical context. The second is a python wrap and release of Brewer's simulated-tempering-like sampler for dealing with multimodal distributions. The third is a short paper that shows that you can figure out the brightness function of stars below the brightness levels at which you can reliably resolve them in a crowded field. The latter project fits into the Camp Hogg brand of measuring the undetectable.

2012-03-23

reverberation mapping

Aaron Barth (UCI) gave our seminar today, about reverberation mapping to get black-hole masses in galaxies with active nuclei. He showed results on attempts to truly invert the problem of the distribution of gas; Brendon Brewer (UCSB), who will be visiting me next week, has been doing some great stuff there with his collaborators at UCSB. Barth emphasized that reverberation mapping is not a precise tool, but it seems to me that with serious modeling of a serious spectroscopic observing program, it perhaps could be.

2012-03-22

after SDSS-III

On the train to and from Baltimore, I worked on typing up everything Fergus has done so far in modeling data from the P1640 spectroscopic coronograph. It is all words, and one heck of a lot of them, but I somehow feel like equations will be even more confusing.

I spent the full work day at JHU in a meeting about the projects that are being assembled to make use of the SDSS 2.5-m Telescope and its spectrographs in the 2014 and beyond period. It was a great meeting, chaired by AS3 (yes "After SDSS-III") Director Mike Blanton (congratulations!). So much was conveyed and discussed in one intense day, I could never properly summarize it. But here are a few highlights for me:

Blanton reviewed the status and budget; with the current imagined scope the budget really can't be less than 40M USD; at full scope and a southern hemisphere companion instrument it gets close to 60M USD. That's not cheap, so we have to create great things. Bundy overviewed the MaNGA project to perform full IFU spectroscopy on 10,000 nearby galaxies to do lots of science and create a massive legacy data set. Drory convinced me that fiber bundles—at least mass-produced fiber bundles— are not a solved problem and represent considerable project risk. During Wake's talk about sample selection a brief battle (started by Drory) broke out about selecting on purely observational properties relative to inferred physical properties. It was a great discussion and hit many philosophical issues of the responsibility for responsible hypothesis testing between observers and theorists and so on. Of course I scare-quoted the above words because the distinction is not totally sound.

Kneib talked about the eBOSS project to continue measuring the baryon acoustic feature at many redshifts with many probes. His thinking about all this has led to a number of enormous proposals (my name on some of them) going in to large telescopes to do enormous imaging projects for selection and weak lensing. I have found a great audience for Holmes, Rix, and my paper on self-calibration. He also showed that he can easily find emission line galaxies at a wide range of redshifts in SDSS and measure their redshifts spectroscopically with the BOSS spectrograph. He showed a galaxy at redshift 1.6 measured with BOSS. Newman showed that WISE does an amazing job of selecting redshift-unity luminous red galaxies. It can select them confidently at magnitudes fainter than we can confidently get spectroscopic reshifts! Crazy! Green overviewed the ambitious TDSS, which plans—I love this— to take a spectrum of every (in some well defined sense) variable object in the PanSTARRS imaging footprint. Awesome!

I got some reactions out of the crowd on three points: I argued that spectrophotometry for the IFU survey is fundamentally different than it was for the previous three SDSS surveys. The arguments for this are subtle and I should write them down carefully. I argued that every MaNGA cartridge should be different by design to maximize efficiency. That was shouted down because of complexity; I don't in fact disagree. I argued that even if you can't get a redshift for a lot of the faint LRGs, they still might be useful for constraining the baryon acoustic feature. That just got laughs!

2012-03-21

why stop at a few eigenvectors?

My loyal reader knows that I hate PCA. That said, Fergus and I are using it to model Oppenheimer's P1640 imaging spectrographic coronograph. We find that the right number of eigenvectors to use from the PCA is in the hundreds not the few to dozen that astronomers are used to. The reason? Fergus and I need to represent the data at exceedingly high accuracy to find faint companions (think exoplanets) among the speckly noise. That said, with hundreds of components, a PCA can properly model any exoplanet and all the speckles, so we use a train and test framework, in which the pixels of interest for finding the exoplanet are not used in building the eigenvectors (which are then subsequently used to model the pixels of interest). That permits us to go to immense model complexity without over-fitting. I love it because it is so crazy; we are barely even compressing the signal with the PCA; we really are just using the PCA to figure out if the pixels of interest are outliers relative to the pixels not of interest. Of course, because all pixels are of interest, we are cycling through all choices of the pixels of interest (and their complementary training set). My job is to write this all up. By Friday! Luckily I am on a train to Baltimore tomorrow. If you have recently sent me email: Expect high latency.

2012-03-20

NASA-Sloan Atlas

Blanton has made the beautiful and useful NASA-Sloan Atlas of galaxies. This is very relevant to my plan for making a human-viewable atlas of the very largest (angularly) galaxies on the sky. Because Blanton was not concerned with precise measurements of the absolutely hugest galaxies, I think we will probably have to use Mykytyn's re-measurements made with the Tractor, but the NASA-Sloan Atlas gives me something pretty good to work with while Mykytyn gets a robust pipeline working. I wrote code today to read, manipulate, and plot the tabular data in Blanton's atlas.