Showing posts with label catalog. Show all posts
Showing posts with label catalog. Show all posts

2024-01-14

divide by your selection function, or multiply by it?

With Kate Storey-Fisher (San Sebastián), Abby Williams (Caltech) is working on a paper about large-angular-scale power, or anisotropy, in the distribution of quasars. It is a great subject; we need to estimate this power in the context of a very non-trivial all-sky selection function. The tradition in cosmology is to divide the data by this selection function. But of course you shouldn't manipulate your data. Instead, you could multiply your model by the selection function. You can guess which one I prefer! In fact you can do either, as long as you weight the data in the right way in the fit. I promised to write up a few words and equations about this for Williams.

2023-07-17

using catalogs responsibly

I had a conversation with Vedant Chandra (Harvard) today about how catalogs are used, and how that relates to how they are built. We started off by arguing about how principled one should be about doing a populations inference. Too abstract! So Chandra moved us in the pragmatic direction: Let's look at a very specific inference and see what matters about it. We decided to look at the distances to distant clusters in the ESA Gaia data: How do your inferences depend on the number of stars you use, the signal-to-noise ratios of those stars, and whether your individual-star measurements are maximum-likelihood or obtained by consideration of a posterior pdf? That should answer questions, and set up some concrete points of discussion.

2023-03-13

something in astronomy is wrong

Today I chatted with my old friend Phil Marshall (SLAC) about various things. Well actually I ranted at him about catalogs and how their use is related to the way they are made. He was sensible in reply. He suggested that we write something, and maybe also develop some guidance for the NSF LSST developer community. His recommendation was to create some examples that are simple but also that connect obviously to research questions in the heads of potential readers. Easier said than done! I said that any such paper needs a good title and abstract. He pointed out that this is true of every paper we ever write! Okay fine.

2023-02-28

purity and completeness

Today Kate Storey-Fisher (NYU) and I discussed how to estimate the stellar contamination of her Gaia and WISE quasar catalog. Because there are few large, complete samples of anything, it’s hard to do this by comparison with any kind of Ground Truth™. What we realized on the call is that it’s easier to estimate how the contamination *changed* as we went from the Gaia quasar candidate table to our final sample. We discussed how to use what external data we have to estimate this.

2023-02-25

catalogs rant

Should I write this paper?

Abstract: Observational astronomy projects often produce catalogs—of stars, galaxies, quasars, planet hosts, and so on—for use in other projects. How can we use these catalogs responsibly? The answer to this turns out to be complex; it depends sensitively on how the catalogs were made. In particular, if the catalog entries were obtained by operations on a set of (nearly) independent or separable likelihood functions, the catalog can be used in a much wider set of circumstances than if the catalog entries were obtained by operations on a posterior pdf or on likelihood functions involving important shared parameters or shared data or shared prior information. This is true no matter whether the subsequent analyses of the catalog are Bayesian or frequentist. Importantly, at the present day, many important catalogs are being made from the outputs of MCMC runs or discriminative machine-learning methods (classifications or regressions). These catalogs are very hard or even impossible to use for population studies. I demonstrate these points mathematically, and also with toy examples from comology, stars, and exoplanets. I recommend that catalogs be designed and made with the feasibility of particular end-user investigations as explicit requirements.

2022-08-14

spherical-harmonic transforms of point sets

On the weekend I computed the spherical-harmonic transform of Kate Storey-Fisher's quasar sample made from ESA Gaia data. I also computed the spherical-harmonic transform of the random catalog we use to map the selection function. The two transforms are extremely similar in their complex amplitudes! Since the random catalog is made assuming perfect homogeneity and isotropy, this similarity directly translates into a measurement of the isotropy of the Universe.

2022-08-04

start a paper on homogeneity

After making (yesterday) all the plots that demonstrate the uniformity and large-scale homogeneity of Kate Storey-Fisher's Gaia quasar catalog (which she is writing up now), I decided (tentatively) to write a paper on cosmic homogeneity with these data: When a catalog shows beautiful homogeneity, that is both a statement about the catalog and a statement about the Universe. I wrote a title and abstract and some figure captions today.

2022-08-01

new paper scope

We had a great meeting today with Kate Storey-Fisher (NYU), Hans-Walter Rix (MPIA), Christina Eilers (MIT), and me to discuss KSF's progress on the ESA Gaia quasar sample. We looked at her large-scale structure results and her jackknifes and discussed paper scope. Options range from a quasar-catalog paper to a selection-function paper to a full cosmological parameter-estimation paper. Of course we decided to do all three! But importantly we decided that this week we would focus on writing a quasar-catalog paper. That's good, and achievable.

2022-02-11

variable-rate Poisson process

At lunch today, Adrian Price-Whelan (Flatiron) challenged me to explain the form of the variable-rate Poisson process likelihood function. I waved my hands! But I think the argument goes like this (don't quote me on this!): You boxelize your space (time or space or whatever you are working in) in small-enough boxels that every boxel contains exactly one or zero data points. Then you take the limit as the boxel sizes go to zero (and become infinitely numerous). The occupied boxels deliver a sum (in the log likelihood) of log densities at the locations of the observed data. The unoccupied boxels deliver an integral of the density over all of space. Something like that?

2021-09-06

a conditional form of domain adaptation?

I have been discussing with Soledad Villar (JHU) projects related to domain adaptation (also with Thabo Samakhoana and Katherine Alsfelder), in which you have two different instruments (say) taking data and you want to find the transformation between the instruments such that the data are the same from the two sources. The modification we want (or need) to make is to create a conditional version of this: In SDSS-V, there are two observatories; they take very similar (but not identical) data. What are the transformations that make the data identical? The problem is: The two observatories also observe different stars on average (because they see different parts of the Galaxy). So we need to find the transformations that make the data identical, conditional on other data (like the ESA Gaia data) that we have for the stars. Great problem, and we came up with some non-elegant solutions. Are there also elegant solutions?

2021-06-28

SDSS-V MWM target selection

Today Jennifer Johnson (OSU) crashed the weekly Manhattan-area SDSS-V discussion meeting to give us the current state of Milky Way Mapper (a component of SDSS-V) target selection. It was a great discussion, because there are many, many target categories, and many of them are interesting to Manhattan-area locals. For me the most impressive thing about the meeting was that Johnson could answer almost any question from anyone on any of the literally dozens of target categories! It was a tour de force as they say. And we learned a lot. One of my goals with this meeting (which was started and is operated by Katie Breivik, Flatiron) is to increase excitement in Manhattan for SDSS-V and Johnson did that admirably.

2021-06-21

SDSS-V meeting

Katie Breivik (Flatiron) has started a meeting in New York for those interested in SDSS-V data and science. This has been fun; I have learned about a lot of different projects that I didn't know about. In today's meeting, Adrian Price-Whelan (Flatiron) showed some plots of the distribution of different abundances in the Milky Way disk, showing that we can probably see the Galaxy mid-plane way better in the abundances than in the kinematics. And kinematic evolution aligns the abundances with the kinematics! Nice result there. We vowed to have an expert come and walk us through SDSS-V target selection soon, since we were all soft on what, exactly, we would target!

2021-04-06

causal-inference issues

I had a nice meeting (in person, gasp!) with Alberto Bolatto (Maryland) about his beautiful results in the EDGE-CALIFA survey of galaxies, and (yes) patches of galaxies. Because they have an IFU, they can look at relationships between gas, dust, composition, temperature, star-formation rate, mean stellar age, and so on, both within and across galaxies. He asked me about some difficult situations in undertanding empirical correlations in a high dimensional space, and (even harder) how to derive causal conclusions. As my loyal reader might guess, I wasn't much help! I handed him a copy of Regression and Other Stories and told him that it's going to get harder before it gets easier! But damn what a beautiful data set.

2021-01-08

Gaia Unlimited kick-off

Today was the first meeting of the Gaia Unlimited project (PI: Anthony Brown), in which we attempt to make a selection function (and the tools for making many different kinds of selection functions) for investigators making use of the ESA Gaia data to perform population-level inferences. Among the many things we discussed were the definition of the selection function (which is not trivial, given the historical usage of the term (and it appears in 700 refereed publications in 2020, according to NASA ADS), and what's known about the Gaia selection function already. The latter includes amazing work by Boubert and Everall in which they have tried to reverse engineer everything to determine the selection function to very faint levels, given the Gaia scan patterns, telemetry limits, and dropped fields. So far, my role in this project is on the conceptual side, around definitions, terminology, and use cases. Along those lines there was great discussion about what the selection function is, and what it is not. Our position is that it is the probability, given hypothetical properties q, that a counter-factual source with those properties would enter the catalog. Even that definition is not quite complete, because there are details relating to the observability of—and noise in—the properties q. More about this throughout this year!

2021-01-06

stream finding

Stars and Exoplanets meeting at Flatiron (well on xoom, really) was all about finding stellar streams. Matt Buckley (Rutgers) talked about repurposing to the astrophysical domain machine-learning methods employed in high-energy physics experimental data to find anomalies. Sarah Pearson (Flatiron) talked about building things that evolve from Hough transforms. In both cases we (the audience) argued that the projects should make catalogs of potential streams with low (non-conservative) thresholds: After all, it is better to find low-mass streams plus some junk than it is to miss them: Every stream is potentially uniquely valuable.

2021-01-04

mathematical derivation of our clustering estimator

In a long conversation, Kate Storey-Fisher (NYU) and I worked through her new and nearly complete derivation of our continuous-function estimator for the correlation function (for large-scale structure). We constructed the estimator heuristically, and demonstrated its correctness somewhat indirectly, so we didn't have a good mathematical derivation per se in the paper. Now we do!

2021-01-03

selection function disagreements

Hans-Walter Rix (MPIA) and I had a solid conversation today about the scope of our first paper on the selection function. We want to be pedagogical in scope and content. So the argument between us is: How sophisticated to get in our selection-function model? Rix is arguing for a less sophisticated case, keeping the story and main point simple, and I am arguing for something more sophisticated, that more connects to the real decisions that people are making every day. And all this relates to exactly what toy problems we show. We came to something of a compromise position, in which we give an example where the apparent magnitude cut is the main selection, but then show what happens when you expand the sample such that other effects beyond the pure apparent magnitude cut start to affect the sample significantly. One of our points will be that as particular selection effects get fractionally smaller in impact on your sample, you don't have to model them as precisely to meet some global accuracy goals for your model of the whole population.

2020-12-31

the selection function; talking it out

I had a quick call today with Hans-Walter Rix in which we had the millionth conversation about the selection function and how it is used in astronomy problems. And how it ought to be used. We resolved our differences, once again. I think it's important to take any project I'm working on and discuss it over and over again. I learned this from various mathematics colleagues, including Goodman (NYU), Greengard (Flatiron), and Villar (JHU). You don't really understand a project until you can rigorously describe it. And you can only rigorously describe it after quite a few trials and errors.

2020-11-24

how many black holes?

I spoke with Katie Breivik (Flatiron) today about a project to paint toy binary stars (from Breivik's model of how binaries form and evolve) onto toy spectroscopic targets (from Neige Frankel's model of how the Milky Way disk formed) to see how many binary stars and how many black-hole (or compact-object) binaries Adrian Price-Whelan (Flatiron) and I should be finding in the APOGEE survey. The project is simple in principle, but the matching up of differently simulated catalogs is a conceptual and administrative challenge! The hope for this project is that we can constrain something about the formation of black-hole binaries by the observation that we don't find any (or don't find very many) in APOGEE!

2020-11-21

selection function and white dwarfs

Hans-Walter Rix (MPIA) and I worked on our project to explain, elucidate, and determine the selection function (for ESA Gaia and other surveys) today. We decided that a good toy problem is the luminosity function of white dwarf stars. This is a good toy problem because the selection function is necessary, but they aren't so far away that the full three-dimensional dust map is required. I wrote words about this problem in a latex document to get us started.