One of the things I have been saying for a few years is that the astronomical images available on the Web, taken together, form some kind of very heterogeneous and odd sky survey. Could this be used to do science? Lang and I have an argument that it can: He Web-searched "comet holmes", took all images found that way, calibrated them with Astrometry.net (many didn't calibrate because, for example, they were images of cats or grandmothers), and fit a gravitational trajectory in the Solar System to the image locations. We started to write this up today.
2009-10-13
2009-10-12
done and done
Astrometry.net paper up on arXiv today; noise-information paper going up tomorrow. My research today was just finishing tweaks on the latter.
2009-10-09
arXiv issues
I spent most of the day chatting, notably with Stumm, who is in NYC for a few days. Late in the day, Lang (attempted to) put the Astrometry.net paper up on the arXiv. This was a failure, and not for the usual reasons: None of the figures got flagged as too big!
The paper compiles without complaint by "pdflatex astrometry-dot-net" on any unix (mac or Linux) platform we have been able to try. And yet on arXiv, the figures can't be understood! (It can't determine size or bounding box or format or even see the file in some cases.)
The arXiv system has a number of issues, not the least of which is that it doesn't just run vanilla pdflatex, ever. It runs in some strange box in which various pdflatex options have been changed from their default values. This would be fine if the arXiv system exposed its pdflatex or latex configuration. But it doesn't. Even better, since all the processing is done by robots, a "sandbox" robot could be established for people to test uploads before submission, greatly reducing the time wasted on this, not to mention the stress; documents and tarballs could be tested as they are being written and not "by fire" right at posting time.
Indeed, the inscrutable arXiv robot will reject a submission on any number of grounds, many of which are mentioned in the arXiv help pages, but few of which are described in enough detail for a user to reliably avoid them. For example, the figure size constraints (which were not our problem today) are never stated explicitly in the help pages; the help pages (like this one) only say that figures "should" be made "small" because that is more "efficient"!
As my loyal reader knows, I love the arXiv; it has transformed astrophysics and all of the sciences. Now lets just make it easier to use! Note that anything that makes it easier to use also makes it easier to maintain and run. (Think of all the emails and blog posts that could be saved!)
2009-10-08
done and ExxonMobil
Lang and I finished the primary Astrometry.net paper for submission! We will put it on the arXiv tomorrow. I am extremely excited to see it go out. Congratulations to Lang and the team.
At the end of the day the Physics Colloquium was by Halsey from ExxonMobil on carbon sequestration. It is amazing that this is on the table, because it is so expensive, it would cost as much to sequester the carbon as it currently costs to produce the oil; that is, it would double the price of oil. The talk reinforced the point that there are no real technological solutions; if we are going to reverse global warming, carbon emissions, and pollution there have to be social, cultural and political changes. Technology can only play a small role.
2009-10-07
frustrating day
I just failed (though only just) to complete the noise-information paper (with Price-Whelan) for submission today. And then I also failed to complete the Astrometry.net paper as well! But both papers are close, so by Friday, if all goes well.
2009-10-06
first draft and last details
I finished the first draft of my contribution to the Blandford meeting at the end of the day today.
At the beginning of the day, Price-Whelan and I worked on last details for his paper on the information in astronomical images.
2009-10-05
unconverged chains
I spent the day talking to anyone who would listen about unconverged MCMC (or equivalent) chains. The issue is a big one, and I have many thoughts, all unorganized. But basically, the point is that when likelihood calls take a long time (think weeks), then there is no way in hell we will ever have a converged and dense sampling of the posterior probability distribution for any parameter space, let alone a large one. At the IPMU meeting last week, most practitioners thought it was impossible to work without a converged chain, but I noted that we never have a converged chain in the larger space of all possible models; whenever we have a converged chain it is just in some extremely limited and constrained subspace (for example the 11-dimensional space of CDM or the like; this is a tiny subspace of all the possible cosmological model spaces). The fact that we don't have a converged chain on all the possible models and all the possible parameters conceivable does not prevent us from doing science. This has connections to the multi-armed bandit problem. I also have been thinking about Rob Fergus's 80 million tiny images project, which treats the result of a huge set of Google (tm) searches as a sampling of the space of all natural images. Of course this is not a converged or dense sampling! But nonetheless, science (and engineering) can be done, very successfully.
2009-10-02
writing
I spent the flight home working on my contribution for the Blandford fest. The SFoA meeting—well, really the side conversations and coffee talk—definitely helped me sharpen up some of the issues.
2009-10-01
SFoA, day 4
Today was exoplanet day at the SFoA meeting, which meant I learned the most. Among other things, Winn (MIT) told us that he can measure the alignments of planetary orbits with stellar rotation and that planet transits more-or-less directly tell you the surface gravity on the planet. Turner (Princeton) spoke about strong biases in fitting planets to radial velocity curves, which emerge from the nonlinearity of the fitting. Laredo (Cornell) showed a system for optimizing or guiding future radial velocity measurements given measurements in the past, where he is optimizing for information gained about the orbit. As he points out, you can optimize for many things there; it is a multi-armed bandit kind of problem. One thing that surprised me is that he takes the convergence of his MCMC chains very seriously, which is good in general, but does not seem necessary to me in order to perform these experimental design activities. After all, your utility will always be approximate, why spend millions of hours of CPU time to fill out in enormous detail predictions of future trajectories that will only be used to approximately calculate your utility? But he is certainly thinking about the problems the right way.
Several mentioned that although orbit fitting is now being performed in quasi-optimal ways, the original data reduction (going from spectral pixels to radial velocity measurements with errors) is not. Turner opined that if this were done better, more planets would be discovered, because there are many detections close to the current limits. That's a big—but very important—problem.
2009-09-30
SFoA, day 3
Today was cosmology day at the meeting, with (among other interesting contributions) a nice discussion by Marinucci of needlets, a class of compact wavelets on the sphere. These have extremely nice properties for statistical analyses of fields on the sphere. Once again I am amazed at all the great things you can do with linear functions of your data! In off time, I spent some time getting very specific advice from Baines about MCMC improvements.
2009-09-29
SFoA, day 2
Today was the second day of the Statistical Frontiers of Astrophysics meeting in Tokyo. Among other very interesting talks, Xiao-Li Meng and Paul Baines worked us through some ideas in the modern use of MCMC. They are particularly interested in problems of data augmentation, where parameters are introduced for every data point (that is, there are more parameters than data points). They had a number of basic ideas for MCMC that are almost always a good idea; I have lots to do to my code when I get home (and lots of new literature to read).
2009-09-28
flight
On my flight to Tokyo (and in an Izakaya here) I worked on various statistics writing projects, including my talk (of course) and various long-term things, each of which is starting to look like a book chapter.
2009-09-25
modeling, modifying
Sam Roweis spoke at the NYU Computer Science Colloquium, and Lam Hui spoke at the NYU Astrophysics Seminar today. Roweis spoke about fitting models to data, where the data are taken by sensors with unknown properties (for example microphones with unknown thresholds, saturation, nonlinearity, and frequency response or sensors on wandering robots of unknown position). His point was that if you have enough sensor readings from differently unreliable or unknown sensors, you can still do very well, if you take a generative modeling approach. The demos were nice. Hui spoke about modifications to gravity and what they might do to dynamics. In particular, he noted that most current modifications to gravity look very much like a scalar–tensor theory, where there is (effectively) a different value of G in high-density regions than in low-density regions. If this is going on in our Universe, there ought to be lots of dynamical signatures; hopefully we can rule out large classes of models of that type.
2009-09-24
HST rules, as usual
The most attentive of readers would know that Price-Whelan and I have been working on assessing the bandwidth necessary for transmitting astronomical image information. That is, we have determined the precision to which you have to record pixel values to preserve all the scientific information in an image, where we assess that information for both bright sources and faint sources below the noise
. Today we pulled down a RAW (unprocessed) HST ACS image from the HST Archive and showed that, in fact, the ACS raw data are telemetered within one bit of the minimum bandwidth. That is, of the 16 bits in the integer representation of raw ACS pixel data, at most one bit is wasted. That's pretty close to optimal, and we will say so in the paper we hope to submit within a week or two.
2009-09-23
stream perturbations
Spent time with Johnston today at Columbia and she encouraged me to dust off my manuscript on perturbations of cold streams by compact substructures. We realized that there are already morphological features in the known streams that could be analyzed, at least roughly, in terms of perturbations by substructures. I started to get her to agree that the smoothness and straightness of streams like GD-1 already make them interesting for the dark-matter model. But we both agreed that we can't be quantitative about that until we understand how the disruption of streams by substructure affects their detectability.