Showing posts with label drinking. Show all posts
Showing posts with label drinking. Show all posts

2022-07-07

is it ever scientifically conservative to use machine learning?

I gave a talk Is machine learning good or bad for science? in Vienna today (slides here). I spent a lot of time on the ontology and epistemology of it all. One thing that led to some debate afterwards is my claim (at the end of the talk) that using extremely flexible machine learning methods can be extremely conservative in some cases: If you are modeling a nuisance that possibly interferes with your signal of interest, and you used a very flexible model, you have a strong argument that you tried as hard as you could (in some sense) to dilute your signal of interest with that nuisance. My talk was followed by interesting discussion with many, and a lovely dinner with Viennese (not just Austrian, but Viennese) wine.

2022-07-06

Dr Ratzenböck

It was my pleasure to be a part of the PhD committee for Sebastian Ratzenböck (Vienna), who wrote a dissertation in computer science but as applied to astrophysics. He had three advisors, in statistics, in computer science, and in astronomy, and he beautifully bridged the three worlds. His research was on finding members of stellar clusters, and on finding new stellar clusters. He showed (pretty convincingly, I think) that star-forming regions break up into many individual star-forming events with different ages and different kinematics. One of his conclusions is that all star formation happens in clusters or groups! He also made a nice technical advance, which was to build a tool to select clustering hyper-parameters in the space of physical quantities one cares about, instead of in the space of arbitrarily-defined clustering-method parameters. It was a great thesis, a beautiful defense, and a fun time drinking afterwards. Congratulations Dr Ratzenböck!

2020-01-07

remote meetings

A highlight of today was a long meeting with Chris Lintott (Oxford) covering many subjects. But he told me about dot-dot-astronomy, which is a fully-remote reboot they are working on for the niche but extremely influential dot-astronomy meetings. The idea is to go fully remote—all participants remote—but then change the meeting expectations and structure to respect that. The idea is: Maybe not try to do remote meetings so they are just as good as face-to-face meetings, but to try to do remote meetings so they are something very different from face-to-face meetings. That seems like a great idea. Let's re-frame our goals. We have to do something about what we are doing to this planet.

2019-05-24

LIGO housekeeping data, MCMC

Yesterday I gave a talk about data science at Oregon Physics. Today I talked about dark matter—on the chalk board. I talked about various bits of vapor-ware that we are doing with ESA Gaia and streams and The Snail. I also talked about the GD-1 perturbation found by Bonaca and Price-Whelan. That was followed by an excellent and fun lunch with the graduate students, in which they interviewed me about my career and science. Oregon Physics has a great PhD cohort.

In the afternoon, Ben Farr (Oregon) and I hived off to a rural brewery to discuss LIGO systematics and MCMC sampling. I have fantasies about calibrating the LIGO strain data using the enormous numbers of housekeeping channels that are recorded simultaneously with the strain. Farr was encouraging, in that he does not believe that big models of this sort have been seriously done inside the Consortium. That means there might be a role for me or for a collaboration that includes me.

On the MCMC front, we discussed a few different sampling ideas. One is a project by Farr called kombine, which is an ensemble sampler that uses the ensemble to inform an approximation to the posterior, which in turn informs the sampling. Another is vapor-ware by me called bento box which hierarchically splits your problem into a tree of disjoint problems until you get to a set of unimodal problems that are individually trivial. I realized while I was talking I could even use HMC with reflection moves to simplify the problem at the hard boundaries of the boxes in the bento box.

On the drive back to the airport, we found that we agreed on the point that no-one should ever compute a fully marginalized likelihood. That was refreshing; Farr is one of the very few Bayesians I know who get this point. It inspired me to spend my time at the airport tinkering with the paper I want to write on this subject.

2019-04-16

binaries

Great Astro Seminar today by Carles Badenes (Pitt), who has been studying binary stars, in the regime that you only have a few radial-velocity measurements. In this regime, you can tell that something is a binary, but you can't tell what its period or velocity amplitude is with any precision (and often almost no precision). He showed results relevant to progenitors of supernovae and other stellar explosions, and also exoplanet populations. Afterwards, Andy Casey (Monash) and I continued the discussion over drinks.

2016-08-07

ready to resubmit

I worked on the weekend to get my “Chemical tagging can work” paper ready for resubmission to the ApJ, incorporating referee and co-author comments, both of which made the paper much better. By Sunday it was good enough to send to the co-authors for final comments. In case it is some comfort to my loyal reader, it took me a full six months to get to this, which is embarrassing, but normal. And even then—when I sent it to the co-authors—it was missing a paragraph about the abundances in cluster M5. While Andy Casey and I were relaxing in a Heidelberg pub, he (Casey) wrote that final paragraph. I love my job!

2014-09-16

AstroData Hack Week, day 2

The day started with Huppenkothen (Amsterdam) and I meeting at a café to discuss what we were going to talk about in the tutorial part of the day. We quickly got derailed to talking about replacing periodograms and auto-correlation functions with Gaussian Processes for finding and measuring quasi-periodic signals in stars and x-ray binaries. We described the simplest possible project and vowed to give it a shot when she arrives at NYU in two months. Immediately following this conversation, we each talked for more than an hour about classical statistics. I focused on the value of standard, frequentist methods for getting fast answers that are reliable, easy to interpret, and well understood. I emphasized the value of having a likelihood function!

In the hack session, I spoke with Eilers (MPIA) and Hennawi (MPIA) about measuring absorption by the intergalactic medium in quasars subject to noisy (and correlated) continuum estimation. Foreman-Mackey explained to me that our failures on K2 the previous night were caused by the inflexibility of the (dumb) PSF model hitting the flexibility of the (totally unconstrained) flat-field. I discussed Gibbs sampling for a simple hierarchical inference with Sick (Queens). And I went through agonizing rounds of good-ideas-turned-bad on classifying pixels in Earth imaging data with Kapadia (Mapbox). On the latter, what is the simplest way to do clustering in the space of pixel histograms?

The research day ended with a discussion of Spectro-Perfectionism (Bolton and Schlegel) with Byler (UW). I told her about the long conversations among Roweis, Bolton, and me many years ago (late 2009) about this. We decided to do a close reading of it (the paper) tomorrow.

2014-09-15

AstroData Hack Week, day 1

On my way to Seattle, I wrote up a two-page document about inferring the velocity distribution when you only get (perhaps noisy, perhaps censored) measurements of v sin i. When I arrived at the AstroData Hack Week, I learned that Foreman-Mackey and Price-Whelan had both come to the same conclusion that this would be a valuable and achievable hack for the week. Price-Whelan and I spent hacking time specifying the project better.

That said, Foreman-Mackey got excited about doing a good job on K2 point-source photometry. We talked out the components of such a model and tried to find the simplest possible version of the project, which Foreman-Mackey wants to approach by building a full, parameterized, physical model of the point-spread function, the spacecraft Euler angles, and the flat-field. Late in the day (at the bar) we found out that our first shot at this model is going badly off the rails: The flat-field and the point-spread function are degenerate (somewhat or totally?) in the naive model we have right now. Simple fixes didn't work.

2014-09-11

single-example learning

I pitched projects to new graduate students in the Physics and Data Science programs today; hopefully some will stick. Late in the day, I took out new Data Science Fellow Brenden Lake (NYU) for a beer, along with Brian McFee (NYU) and Foreman-Mackey. We discussed many things, but we were blown away by Lake's experiments on single-instance learning: Can a machine learn to identify or generate a class of objects from seeing only a single example? Humans are great at this but machines are not. He showed us comparisons between his best machines and experimental subjects found on the Mechanical Turk. His machines don't do badly!