Showing posts with label frequentism. Show all posts
Showing posts with label frequentism. Show all posts

2025-07-28

integrating out nuisances

Further insipired by yesterday's post about binary fitting, I worked today on the treatment of nuisance parameters that have known distributions. These can be treated as noise sometimes. Let me explain:

If I had to cartoon inference (or measurement) in the face of nuisance parameters, I would say that frequentists profile (optimize) over the nuisances and Bayesians marginalize (integrate) over the nuisances. In general frequentists cannot integrate over anything, because there is no measure in any of the parameter spaces. But sometimes there is a measure. In particular, when there is a compact symmetry:

We know (or very strongly believe) that all possible orientations of a binary-star orbit are equally likely. In this model (or under this normal assumption) we have a distribution over two angles (theta and phi for that orbit pole, say); it is the distribution set by the compact group SO(2). Thus we can treat the orientation as a noise source with known distribution and integrate over it, just like we would any other noise source. So, in this case (and many cases like it) we can integrate (marginalize) even as frequentists. That is, there are frequentism-safe marginalizations possible in binary-star orbit fitting. This should drop the 12-parameter fits (for ESA Gaia data) down to 8-parameter, if I have done my math right.

2025-07-24

how significant is your anomaly?

So imagine that you have a unique data set Y, and in that data set Y you measure a bunch of parameters θ by a bunch of different methods. Then you find, in your favorite analysis, your estimate of one particular parameter is way out of line: All of physics must be wrong! How do you figure out the significance of your result?

If you only ever have data Y, you can't answer this question very satisfactorily: You searched Y for an anomaly, and now you want to test the significance. That's why so many a posteriori anomaly results end up going away: That search probably tested way more hypotheses than you think it did, so any significances should be reduced accordingly.

The best approach is to use only part of your data (somehow) to search, and then use a found anomaly to propose a hypothesis test, and then test that test in the held-out or new data. But that often isn't possible, or it is already too late. But if you can do this, then there is usually a likelihood ratio that is decisive about the significance of the anomaly!

I discussed all these issues today with Kate Storey-Fisher (Stanford) and Abby Williams (Chicago) today, as we are trying to finish a paper on the anomalous amplitude of the kinematic dipole in quasar samples.

2025-07-21

wrote like the wind; frequentist vs Bayes on sparsity

My goal this year in Heidelberg is to move forward all writing projects. I didn't really want to start new projects, but of course I can't help myself, hence the previous post. But today I crushed the writing: I wrote four pages in the book that Rix (MPIA) wants me to write, and I got more than halfway done with a Templeton Foundation pre-proposal that I'm thinking about, and I partially wrote up the method of the robust dimensionality reduction that I was working on over the weekend. So it was a good day.

That said, I don't think that the iteratively reweighted least squares implementation that I am using in my dimensionality reduction has a good probabilistic interpretation. That is, it can't be described in terms of a likelihood function. This is related to the fact that frequentist methods that enforce sparsity (like L1 regularization) don't look anything like Bayesian methods that encourage sparsity (like massed priors). I don't know how to present these issues in any paper I try to write.

2022-09-28

can a (say) 5-sigma result be interpreted as a p-value?

In preparing for class today (I am teaching NYU #data4physics), I worked through the relationship between a p-value (like what's used in medical research) and a physicist's n-sigma measurement. They are related in some very special cases, like in particular when the value being measured is a linear parameter (like an amplitude) and the noise is Gaussian. But those cases are special. And also: Converting n-sigma to p-value depends very critically on the noise model. So I don't like thinking of it as a p-value. That said, maybe there is no difference?

2019-12-05

new methods for cosmology

Kate Storey-Fisher (NYU) and I had a lunch-time conversation about her project to replace the standard large-scale-structure correlation-function estimator with a new estimator that doesn't require binning the galaxy pairs into separation bins. It estimates continuous functions. We discussed how to present such a new idea to a community that has been using the same binned estimator since the 1990s, and even before that they were only marginally different. That is, the change Storey-Fisher proposes is the biggest change to correlation-function estimation since it all started, in my (somewhat not humble) opinion.

But this creates a problem: How to convince the cosmologists that they need to learn new tricks? We have many arguments, but which one is strongest? We no longer need to bin, and binning is sinning! Or: We can capture more functional variation with fewer degrees of freedom, so we reduce simulation requirements! Or: We can restrict the function space to smooth functions, so we regularize away unphysical high-frequency components! Or: We get smaller uncertainties on the clustering at every scale! Or: We can make our continuous function components be similar to derivatives of the correlation function with respect to cosmological parameters and therefore create clustering statistics that are close to Fisher-optimal given the data and the model!

Writing methodological papers is not easy.