Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

2025-07-24

how significant is your anomaly?

So imagine that you have a unique data set Y, and in that data set Y you measure a bunch of parameters θ by a bunch of different methods. Then you find, in your favorite analysis, your estimate of one particular parameter is way out of line: All of physics must be wrong! How do you figure out the significance of your result?

If you only ever have data Y, you can't answer this question very satisfactorily: You searched Y for an anomaly, and now you want to test the significance. That's why so many a posteriori anomaly results end up going away: That search probably tested way more hypotheses than you think it did, so any significances should be reduced accordingly.

The best approach is to use only part of your data (somehow) to search, and then use a found anomaly to propose a hypothesis test, and then test that test in the held-out or new data. But that often isn't possible, or it is already too late. But if you can do this, then there is usually a likelihood ratio that is decisive about the significance of the anomaly!

I discussed all these issues today with Kate Storey-Fisher (Stanford) and Abby Williams (Chicago) today, as we are trying to finish a paper on the anomalous amplitude of the kinematic dipole in quasar samples.

2024-03-08

combining spectral exposures

I wrote words! I got back to actually doing research this week, in part inspired by a conversation with my very good friend Greg McDonald (Rum & Code). I worked on the words in the paper I am finishing with Andy Casey (Monash) about how to combine individual-visit exposures into a mean spectrum. The biggest writing job I did today was the part of the paper called “implementation notes”, which talks about how to actually implement the math on a finite computer.

2024-01-19

Happy birthday, Rix

Today was an all-day event at MPIA to celebrate the 60th birthday (and 25th year as Director) of Hans-Walter Rix (MPIA). There were many remarkable presentations and stories; he has left a trail of goodwill wherever he has gone! I decided to use the opportunity to talk about measurement, which is something that Rix and I have discussed for the last 18 years. My slides are here.

I've been very lucky with the opportunities I've had to work with wonderful people.

2023-11-27

Terra Hunting Fall Science Meeting, day 1

Today was the first day of the Terra Hunting annual science meeting. One highlight of the day was a presentation by Yan Liang (Princeton), who is modeling stellar spectral variability (the tiny variability) that affects extremely precise radial-velocity measurements. Her method involves a neural network, which is trained to distinguish RV variations and spectral shape variations through a self-supervised approach (with a data augmentation). Then it separates true stellar RV variations from spectral-variability-induced wrong RV variations by requiring (essentially) that the RV variations be uncorrelated with the (latent) description of the stellar spectral shape. This connects to various themes I am interested in, including wobble by Bedell, a spectral variability project by Zhao, and causal structure in machine learning.

2023-11-14

conjectures about pre-training

On Monday of this week, Shirley Ho (Flatiron) gave a talk at NYU in which she mentioned the unreasonable effectiveness of pre-training a neural network: If, before you train your network on your real (expensive, small) training data, you train it on a lot of (cheap, approximate) pre-training data, you get better overall performance. Why? Ho discussed this in the context of PDE emulation: She pre-trains with cheap PDEs and then trains on expensive PDEs and she gets way better performance than she does if she just trains on the expsensive stuff.

Why does this work? One interesting observation is that even pre-training on cat videos helps with the final training! Ho's belief is that the pre-training gets the network understanding time continuity and other smoothness kinds of things. My conjecture is that the pre-training teaches the network about (approximate) diffeomorphism invariance (coordinate freedom). The cool thing is that these conjectures could be tested with interventions!

2023-11-13

radical papers I want to write (or will never write)

I have to finish my NSF proposal with Mike Blanton (NYU), so naturally I am in procrastination mode. Here are three papers I wish I would write. Maybe I should post them on my ideas blog:

Occam's Razor is wrong: This paper, co-authored with Jennifer Hill (NYU), would be about the fact that, in the real, observed world, the simplest explanation is always wrong or at least incomplete.

Causation is just causality: This paper, maybe co-authored with David Blei (Columbia) or Bernhard Schölkopf (MPI-IS) or Hill, shows that you don't need to have free will in order to have cogent causal explanations of data. That is, you don't need to phrase causality in terms of predictions for counter-factual experiments that you might have chosen to do.

You don't ever want evidence: This paper shows that any time you are computing the Bayesian evidence—what I call the fully marginalized likelihood (fml)—you are doing the wrong integral and solving the wrong problem. For both practical and theoretical (principled) reasons.

2023-10-19

Florida, day one

I spent today with Sarah Ballard's group, plus others, at the University of Florida. I gave a talk, to a large, lively, and delightful audience. At the end of this talk I was very impressed by the following thing: Ballard had everyone in the room discuss with their neighbors (turn and talk) for about 3 minutes, after the seminar but before the question period began! This is a technique I use in class sometimes; it increases participation. After those 3 minutes, audience members had myriad questions, as one might imagine.

I spoke with many people in the Department about their projects. One highlight was Jason Dittman, who showed me gorgeous evidence that a particular warm exoplanet on an eccentric orbit has an atmosphere that undergoes some kind of phase change at some critical insolation, as it moves away from its host star on its orbit. Crazy!

Late in the day I discussed n-point functions and other cosmological statistics with Zach Slepian and Jiamin Hou. We discussed the plausibility of getting tractable likelihoods for any n-point functions. We also discussed the oddity that n-point functions involve sums over n-star configurations among N stars (N choose n), but there are mathematical results that show that any permutation-invariant function of any point cloud can be expressed with only a sum over stars (N). That sounds like a research problem!

2023-10-18

biases from machine learning

Today I gave a talk (with these slides) at a meeting in Denver for the NSF initiative Harnessing the Data Revolution. I spoke about the necessity and also the dangers of using machine-learning methods in scientific projects. I brought up two very serious possible biases. The first is that if emulators are used to replace simulations, and they can't be easily checked (because the simulation requirements are too expensive), the emulators will lead to a confirmation-bias problem: We will only carefully check the emulations if they lead to results that we don't like! The second bias I raised is that if we perform joint analyses on objects (stars, say) that have been labeled (with ages, say) by a machine-learning regression, there will in general be strong biases in those joint analyses. For example, the average value of 1000 age labels for stars labeled by a standard ML regression will not be anything like an unbiased estimate of the true average age of those stars. These biases are very strong and bad! That said, I also gave many example locations where using machine learning methods is not just okay but actually intellectually correct, in areas of instrument calibration, foregrounds, and other confounders.

The question period was great! We had 25 minutes of questions and answers, which ranged across a very wide set of topics, including statistics, experimental design, and epistemology.

2023-10-17

Bayesian evidence?

Kate Storey-Fisher, Abby Williams, and I spent some time discussing unpublished work that relies heavily on calculations of the Bayesian evidence. Bayesian evidence—what I call the “fully marginalized likelihood”—relates to the volume of the posterior in parameter space. It is generally extremely sensitive to the width of the prior pdf, since if you are comparing two models with different parameterizations, the numbers you get depend on how you normalize or scale out the units of those parameter-space volumes. Indeed, you can get any evidence ratios you want by tuning prior pdf widths. That's bad if you are trying to conclude something, scientifically! Bayesian inference is only principled, imho, when you can quantitatively state the prior pdf that correctly describes your beliefs, prior to seeing the new data. And even then, your evidence is special to you; any other scientist has to recompute from scratch.

2023-09-06

is a periodic signal in a time series statistically significant?

I had conversations with Nora Eisner (Flatiron) and Abby Shaum (CUNY) today about how we report the significance of a signal we find in a time series. In particular a periodic signal. It's an old, unsolved problem, with a lot of literature. And various hacks that are popular in the exoplanet community (and binary-star community!). My position is very simple: Since all methods for determining significance are flawed, and since when you fit a signal you have to estimate also an uncertainty on that signal's parameters, the simplest and most basic test of significance is the significance with which you measure the amplitude of the proposed signal. That is, if the amplitude is well measured, the signal is real. Of course there are adversarial data sets I can make where this isn't true! But that's just a restatement of the point that this is an unsolved problem. For deep reasons!

2023-09-05

teeny tiny cosmological simulations.

Connor Hainje (NYU) is looking at this paper by Chen et al which uses a machine-learning regression to interpolate between cosmological simulation outputs at different cosmological epochs. To build an end-to-end pipeline for testing ideas, he has been running 32-cubed cosmological simulations. These might be the smallest simulations run since the 1980s! But, interestingly, he is finding that the interpolation isn't working great. Is this because it is harder to train a regression on a small simulation than it is on a large simulation? Is a small simulation less predictable or less interpolate-able? It's expensive to find out!

2023-08-31

O-minus-C inanity

In the exoplanet (and, before that, eclipsing-binary) communities, transit-timing variations are described in terms of a quantity called O−C (pronounced “oh minus sea”), which is the difference between the observed transit time and the “computed” transit time. Right now, Abby Shaum (CUNY) and I are using this terminology in our manuscript about phase variations in coherent pulsators with companions, at the behest of Keaton Bell (CUNY). Okay fine! But O−C has this terrible property, which is that the C part depends on the period or frequency you assume. You can completely change the appearance or morphology of an O−C plot just by slightly tweaking the period. And there is no true period of course! There is just whatever estimates you can make. Which are, in turn, affected by what you use to model the O−C. So it is absolutely awful in every way. Not a stable observable, people! Not even identifiable.

2023-08-18

CZS Summer School, day 5: diffusion

Diffusion models are all the rage in machine learning these days. Today Laurence Levasseur (Montréal) gave a beautiful talk at the CZS Summer School about how diffusion works. She started with a long physics introduction, which was great, and also insightful, about how diffusion works in small physical systems. Then she showed how it can be turned into a method for sampling very difficult probability distributions.

I have a history of working on MCMC methods. These permit you to sample a posterior pdf when you only know a function f that is related to your posterior pdf by some unknown normalization constant. Similarly, diffusion lets you sample from a pdf when you only know the gradient of f. Again, you don't need the normalization. That makes me wonder: Should we be using diffusion in places where we currently use MCMC? I bet the answer is yes, for at least some problems.

2023-07-30

building trust in emulators

I started writing in a possible grant proposal (that would be in collaboration with others) about the trustworthiness of machine-learning emulators. Emulators are systems that learn the input–output relationship of a computationally expensive simulation and produce (or speed the computation of) new simulation outputs, reducing total computational requirements for a given number of simulations. These are so important now that the ESA Euclid and Simons Observatory data-analysis plans crucially involve emulation.

The issue is: How do we trust that the emulators are giving good outputs? There is no obvious way to test them, except by comparing to held-out training data. But in large-scale structure contexts, no amount of held-out data can test the enormous input data space. I don't know how we will ever trust such systems (and damn do we need to!), but I have some ideas about how to improve the situation. One involves enforcing physics symmetries on the emulators. Another involves running adversarial attacks on them.

2023-07-21

an insight about machine learning

Gaby Contardo (SISSA) completed her visit to Heidelberg today. Over coffee this morning she delivered a very simple, but very nice insight about machine learning outputs. Apologies that this is very Inside Baseball:

As I like to emphasize, you can't really average (or do any populations inferences with) a collection of labels delivered by a discriminative ML method run on a collection of objects. Think: Finding the mean age of a cluster, where each star in the cluster got an age estimate from a discriminative ML method trained on stars with known ages. This is because the discriminative ML methods output something very akin to posterior quantities, and if you average a bunch of posterior estimates, you are multiplying in a prior times itself many times; eventually the prior dominates the inference (in many cases).

Contardo's point: If what you want is a label for a collection of objects, like that mean age, you should train on collections of objects. That is, make a training set where you have sets of N stars, labeled by mean age. Then this model can be applied to a new collection of stars and deliver a mean age estimate! Haha, brilliant. And correct. And consistent with the rules of inference.

2023-07-18

a likelihood for our Phi-M radio

There are AM radios and FM radios and (if you are a nerd) PCM radios. But Abby Shaum (CUNY) and I have built a Phi-M radio, which demodulates phase variations in a carrier signal. We (with Keaton Bell, CUNY) are using it to find binary companions and planets around stars that show coherent pulsation modes in their photometry. Today I wrote down a noise model for the output of our demodulator. It isn't completely trivial. But it's good, because we can make a likelihood function for fitting our companions. Our model will end up being a limit of the more general model called Maelstrom by Dan Hey (Hawai'i).

2023-07-17

using catalogs responsibly

I had a conversation with Vedant Chandra (Harvard) today about how catalogs are used, and how that relates to how they are built. We started off by arguing about how principled one should be about doing a populations inference. Too abstract! So Chandra moved us in the pragmatic direction: Let's look at a very specific inference and see what matters about it. We decided to look at the distances to distant clusters in the ESA Gaia data: How do your inferences depend on the number of stars you use, the signal-to-noise ratios of those stars, and whether your individual-star measurements are maximum-likelihood or obtained by consideration of a posterior pdf? That should answer questions, and set up some concrete points of discussion.

2023-07-14

data-driven information

My day started with a conversation with Wolfgang Brandner (MPIA), who asked me how to figure out the information content of ESA Gaia RVS spectra, but in a data-driven way. He wants to avoid the theoretical models at first; that is, he wants to figure out how precisely the spectra contain temperature and metallicity and age information without having temperatures, metallicities, and ages that we believe. One approach is to compare to other data that are sensitve to temperature, metallicity, and age: If the RVS spectra can predict those data, then (conditioned on assumptions) they must contain information about temperature, metallicity, and age. This is similar to questions of risk (or expected error in prediction) in machine-learning contexts.

2023-05-31

Dr Kate Storey-Fisher

Kate Storey-Fisher (NYU) defended her PhD here at NYU today. She killed it! She talked about emulating cosmological simulations (at the level of statistics, not maps), making invariant scalars that encode the shapes and dynamics of dark-matter halos, and her awesome 1.2 million all-sky quasar catalog from ESA Gaia and NASA WISE. It was all things my loyal reader knows lots about but I loved it. It has been an honor and a privilege to work with KSF these years, and I will miss her very very much.

2023-05-30

Dr Irina Espejo

Today it was my honor to serve on the PhD defense committee of Irina Espejo (NYU), who is one of the first (ever in the world, actually!) PhDs in Data Science. Her PhD research involved making real, practical, scalable, reproducible tools for the (late-in-pipeline) analysis of high-energy physics data from the Large Hadron Collider. She built tools to speed up likelihood-free inferences, and she built a tool to find exclusion regions (upper limits) in complex parameter spaces. She used the latter to put constraints on a (real, not toy) proposed modification to the standard model.

On the first project, the tools that she built (and built on) make the LHC more sensitive to new physics, because they find better test statistics for distinguishing models. They make some searches far better, which makes me wonder whether particle physics is using our money efficiently??