Thursday, July 9, 2009

Countability and Strong Positive Anymore

Seven or eight years ago, I discovered that I have a linguistic condition called "Positive Anymore." I was in college, chatting with my roommates, and said something like "Anymore, I just lift weights in the Leverett gym." One of my roommates just couldn't take it anymore. "I've heard you say 'anymore' like that for years now. What the hell does it mean?" A quick poll of those present revealed that I was the only one for whom that construction made grammatical sense. My aunt (a linguistics professor) gave me the prognosis: I had Positive Anymore.

I used to think that PA was a linguistic shortcoming of mine, but anymore I'm convinced it's more like a superpower. Whereas most English speakers can only use the word in negative constructions like "I don't drive anymore" I have the uncanny ability to use it positively as per the first sentence in this paragraph. Moreover, I don't just have PA, I have strong PA, which means I can, at will, detach "anymore" and put it anywhere in a sentence. "Anymore, I just take the bus." Astounded yet?
If you're still having trouble parsing that, replace "anymore" with "nowadays"--to me, they mean about the same thing.

I receive no end of flak from friends, relatives, and spouses about my PA, although it's not that uncommon a condition. In fact, I have caught several of my relatives (mostly on my dad's side) using PA even after making fun of me for my PA. They have it and don't even know it.

Anyway, S and I were trying to figure out if my use of PA was inconsistent with my use of "any", which I use according to the standard rules. But it turns out the standard rules are weirder than you might think. For example, "I don't want any spam" is a grammatical negative construction; "I want any spam" is ungrammatical and positive. "Do you want any spam?" is grammatical and seems positive, but it turns out that there is an implied negative in English owing to the uncertainty inherent in questions and subjunctives. Hmph. And then what about "I like any spam I can get?" That seems positive again, but the clause "I can get" is required to make it grammatical. And then the plot thickens. "I feed spam to any dog" is grammatical and did not require a clause to modify "any dog."

Our theory was that using "any" positively without a clause requires the noun modified by "any" to come in quantized units--to be countable. Dogs are countable; spam is a continuum, much like water or space-time. The best example we could think of that illustrated this was "fish". Fish can be countable noun (number of live fishes) or continuum noun (amount of dead fish to eat). If I say "I'll take any fish," the ambiguity in the countability of fish is broken--it's clear I'm talking about a live fish (or, perhaps, a type of fish, which is also countable). But if I say "I'll take any fish that you give me," the ambiguity is preserved.

The implication was that for my use of PA to be consistent with the standard use of "any" (which it may not need to be, since "anymore" is one word, not "any more"), I must be thinking of the time interval referred to by PA (i.e. now and continuing indefinitely into the future) as something countable rather than continuous. Maybe. I don't know. Anymore, I'm just really confused.

Thursday, July 2, 2009

The Need for Speed


On the drive from San Juan to Arecibo this morning, I got to wondering about where my average driving speed fell in the distribution of drivers here in Puerto Rico. In the states, I felt like I was a pretty average driver, but here en la isla, the distribution of driving speeds is different. There are a lot of fast drivers, too be sure, but there is also a subpopulation of drivers whose speed is significantly (~10 mph) below the speed limit. This may be because relative to the US, PR is economically depressed and so more old cars are on the road, or as a reaction to the more erratic driving habits there seem to be here, but anyway, I definitely pass more people than pass me now.

So in an effort to discover where my driving speed fell relative to others (and in and effort to alleviate the boredom of driving 1.5 hrs alone), I started counting how many cars I passed and how many passed me as I was going 65 mph (the speed limit). Out of 55 pass events, only 9 involved me getting passed. To make this a tractable problem in my head, I decided to assume that driving speeds were normally (gaussian) distributed about a mean--even though this contradicts my anecdotal evidence above. Using this approximation, my first instinct was to say that 1/6 of the cars on the road were faster than me, and since ~2/3 of samples are within +/- 1 sigma of the mean in a gaussian distribution, 1/6 of the samples would be above +1 sigma. So I approximated that I was a 1 sigma driver.

But then it occurred to me that I needed to control for a significant sample bias. This is because the test I was doing wasn't randomly selecting cars and comparing my speed to them. Cars were far more likely to get selected if the difference between their speed and my speed was large. A car going the same speed as me would never pass me, and I would never pass it. But I would assuredly pass almost every car on the road that was going 10 mph as I went 65. The "road distance" that I sampled for different velocities is proportional to abs(v-v0), where v0 is my velocity. The effect this had on my samples was to underweight speeds close to my own and overweight the wings of the gaussian distribution I had assumed as my model. If I drove exactly the mean velocity, this effect would not be terribly important--if the model were correct, it would still be the case that as many cars passed me as I passed. But as my velocity moves away from the mean velocity, the "normal drivers" who are only going a little faster than me get undersampled, so I only see the drivers who are tearing around like a bat out of hell. At the lower end, I still see the real slow-pokes on the road, but I start seeing people who are going a bit faster than that, of which there are a lot more. The effect of this sample bias, it seems, would be to make it seem that I'm a farther outlier in my driving speed than I actually am.

So now, to figure out where I fall in the (normal) distribution of driving speeds, I need to know exactly what the mean driving speed and what sigma is, so that I can compensate for the abs(v-v0)
sampling factor. That means I need to figure out 2 numbers, but unfortunately, I only measured 1 number (that 1/6 of the passes while driving were me being passed) so I won't be able to properly constrain this problem. However, I should be able to figure out the mean on my drive home by finding the speed at which as many people pass me as I pass. For now, let's say this is 55 mph. Then all I need to do is find the sigma for which a gaussian distribution around 55 mph downweighted by abs(v-65 mph) has 1/6 of the area lying above 65 mph. I just solved that numerically on my computer, and it's saying that the best-fit sigma is ~22 mph. So that puts me at about +1/2 sigma. That seems reasonable.

An interesting next step (which I'm not going to do right now since I need to get to work) would be to translate the sample error in my pass measurements into an error in the determination of sigma, and then the error in my driving speed percentile.

Tuesday, June 2, 2009

What are the chances?

I cringe every time that I hear this phrase. I heard it most recently when my friend had her car stolen from San Juan. It was recovered in a semi-drivable state in Bayamon. She invested a couple thousand dollars to get rid of the "semi", only to have the car re-stolen a couple of months later. This time, when the car was recovered in Dorado, there was no "semi" to be had, so she's currently trying to sell it for pieces. While the police were fairly understanding (though less than helpful) the first time her car got stolen, the second time was occasion for all sorts of raised eyebrows and skepticism. And in exasperation my friend uttered the phrase in question.

"What are the chances" is a Pandora's box of bad statistics. Statistics is about hedging your bets given incomplete information, but this phrase is always uttered after the fact, when we have (relatively) complete information: it happened. So unless you plan on repeating the experiment, the chances are one. It happened.

As an example, let's take the famous Goat/Car Puzzle. There are 3 doors; one has a car behind it and the other two have goats. After you pick a door, the game-show host opens one of the other doors and reveals a goat. You are then offered the option of switching your choice to the other door. If you play this game repeatedly, you'll win more often if you switch your choice. But the instance just played out, the car was behind one of the doors, and if that was the door you picked, your chances of getting the car were 1. If you didn't, your chances were 0.

You might object: "What were the chances beforehand, when I didn't know where the goat and car were?". But to do that, you need to make some assumptions. You need to assume that at each playing of the game, the cars and goats are randomly assigned and/or you randomly pick doors. Otherwise, it might be the case that the car is always behind door #1, and you always pick door #2. Your chances of success wouldn't be so good in this case. You might have decent prior knowledge of how cars, goats, and doors are picked in this example, but for everyday occurrences, we usually have much more limited prior knowledge. Are cars randomly stolen, or are certain brands targeted? Are certain areas targeted? People often assume that these processes are random, but they rarely are. With limited priors on these events, the question "what are the chances" can't be answered with any certainty and any answers given should be taken with a great big shaker of salt.

Furthermore, people have selective attention. We ignore whole heaps of ordinary outcomes and only pay attention to ones that strike us as interesting. As a friend of mine once said: "Low probability stuff happens pretty regularly because stuff is happening all the time." Even if the processes involved are random, unlikely outcomes are to be expected if the processes are repeated often enough. People tend to ignore the ordinary outcomes, exclaim at the extraordinary ones, and then assume that something deeper is afoot. In my friend's case, the police started wondering if she was being personally targeted or if she was really bad at locking her car. But even if we assume a random model of car thefts, some number unlikely outcomes doesn't automatically imply that our random model is wrong.

Finally, we also need to keep in mind that in complex systems like real life, there may be a huge number of possible outcomes. But something has to happen. When you roll a die, each number only has a 1/6 chance of coming up. Would you roll a die once and then exclaim: "Wow, it came up six! What are the chances?" In real life, there might be millions of outcomes, each with one-in-a-million chance of coming true, but the fact that one of them happens shouldn't be surprising.

Fighting against all of the pitfalls inherent in asking "What are the chances?", I've developed a reflexive response: "What are the chances?" One.

Thursday, May 21, 2009

Ida y Vuelta: a Tail?

It was all over the news and the net yesterday: the missing link has been found! Darwin has finally been proven right!

Hold on a minute. There's no "missing link"--we have example after example of the evolutionary forebearers of Homo sapiens. Evolution was not in question--we already knew Darwin was right. Ida (the name of the fossil found) is not a direct ancestor of humans--the fossil represents a transitional form between lemurs and other apes.

I do not mean to denigrate what is obviously a very important and exceptionally well-preserved fossil that may provide important insights into primate evolution. My objection is to the media frenzy that gives the impression that this is the final resolution to a scientific debate about evolution that a) never existed and b) would not have been resolved by the fossil in question if it did exist.

We have Neanderthals, Homo erectus, Homo habilis, Australopithecus afarensis, and more. Why is Ida, who isn't even our evolutionary forebearer, the "missing link"? I have a theory: it's about the tail. I think that somehow, through all the monkey-human-common-ancestor debate, through the discovery of early hominids, through all of the discussion about increasing cranial capacity and tool usage, everyone has really been wondering "what about the tail? What happened to the tail?" The media was holding out for something it could tout as a human ancestor (or something close enough) with a tail to finally declare the "missing link" found and the evolution debate settled.

You might say that this is all for the best--that the false debate, however belatedly, is finally being put to rest with one last hurrah. But here's a scenario that makes be cringe just to think about it: suppose it turns out that this fossil, which was recovered from a private collection, turns out to be something other than the specimen that the scientist in question thinks it is. Scientists makes mistakes--that's what peer review is for. Then suddenly we've breathed new life into a misconception about evolution that has been hanging around for way too long already.

Tuesday, May 12, 2009

Golomb Rulers/Squares/Rectangles

Yesterday I came across an interesting example of the isolation of academic fields from one another.

A common design parameter for antenna arrays is to try to obtain uniform coverage of the aperture plane to get as many independent measurements as possible. Interferometers sample the aperture plane at locations that correspond to the difference vectors between antenna elements. For example, in one dimension, if I put four antennas at locations (0,1,4,6), then that array would sample the difference set of those positions: (1,2,3,4,5,6). However, if I put antennas at (0,1,2,3), the difference set would only include (1,2,3) with 1 occuring 3 times and 2 occuring twice. For the purpose of uniformly sampling an aperture, redundant spacings are lost measurements. We're looking for a minimum-redundancy array.
There have been a few papers in radio astronomy on minimum-redundancy arrays for the more useful 2-dimensional case, including Golay (1971), Klemperer (1974), and Cornwell (1988).

Thinking that this might be a mathematical problem of interest, I ran it by a good friend of mine: Phil Matchett Wood--a mathematician at Princeton. He quickly uncovered the equivalent problem as formulated in math literature: Golomb rulers in 1-D and Golomb rectangles in 2-D. Some relevant papers on the subject are Shearer (1995), Meyer & Jaumard (2005), Robinson (1985), and Robinson (1997). These papers are on the exact problem and describe applications to "radar and sonar signal detection", obviously referring to the need for independent aperture samples in radar and sonar interferometers. Somehow, the differing nomenclature between these fields was never quite bridged, and so there has not been and cross-referencing between these two formulations of the same underlying problem. People like to talk about the possibility that relevant research in one field goes unnoticed by other fields. This is the first time I've across it myself, though.

Tuesday, May 5, 2009

Compressed Sensing and Wiener Filtering

Today I'm trying to expand my understanding of how we can best remove contaminant signals from the data we take with the Precision Array for Probing the Epoch of Reionization (PAPER). There is a specific problem I want to make sure we can solve for PAPER. Foregrounds to our signal, particularly synchrotron radiation, are expected to be very smooth with frequency. The idea put forth by the MWA and LOFAR groups is that by observing the same spatial harmonics at multiple frequencies, we should be able to remove such smooth components to suppress them relative to the cosmic reionization signal we are looking for. However, generating overlapping coverage of spatial harmonics as a function of frequency is expensive. My intuition is that since foregrounds do not have a spatial structure that changes dramatically with frequency, we shouldn't need to sample a given spatial harmonic very finely in frequency to get the suppression we want. This would allow us to spread our antennas out a little more and get measurements of the sky at a variety of spatial modes.

In many ways, our problem is analogous to what was done with the Cosmic Microwave Background (CMB). For foreground removal in CMB work, Tegmark and Efstathiou (1996) begin with an assumption that foregrounds can be described as the product of a spatial term and a spectral frequency term. This allows them to construct Wiener filters that use the internal degrees of freedom of their data, together with a model of their foreground and a weighting factor based on the noisiness of their data, to construct a filter for removing that foreground. For the most part, this is standard Wiener filtering, except they have to be careful about what they do to their power spectrum, so they apply a normalization factor to correct for a deficiency in Wiener filters. Tegmark (1998) goes on to generalize this technique for foregrounds that vary slowly with frequency. I'm in the process of wading through these papers, but they seem to be directly applicable to what we are doing, and seem to confirm my suspicions that synchrotron emission should be well-enough behaved to require only sparse frequency coverage of a wavemode in order to be suppressed.

Another tactic that I am investigating is that of compressed sensing which I was alerted to in talks by Scaife and Schwardt at the SKA Imaging Workshop in Socorro this last April. The landmark paper on this principle seems to be Donoho (2006), where it is shown that the compressibility of a signal (being sparse for some choice of coordinates) is a sufficient regularization criterion to faithfully reconstruct signals using a small number of samples. In a way, this technique has an element of Occam's Razor in it--it tries to find a solution, in some optimal basis, that needs the fewest non-zero numbers to agree with the measured data. At least, that's my take on it without having finished the paper.

The relevance of compressed sensing to image deconvolution is explored in Wiaux et al (2009), and it seems to be powerful. I'm excited by this deconvolution approach because it meshes well with the intuitive approach I've been taking to deconvolution, which was to use wavelets and a Markov Chain Monte Carlo optimizer to find the model with the fewest number of components that reproduces our data to within the noise. Compressed sensing seems to be exactly this idea, but is agnostic about the basis chosen, instead of mandating one like wavelets. Anyway, this technique may also be relevant to our foreground removal problem because we might be able to use it to construct the minimal foreground model implied by our data. For synchrotron emission, which should have smoothly varying spatial structure with frequency, I envision that this could construct a maximally smooth model that would allow us to use sparse frequency coverage to remove the foreground emission to the extent that it is possible to do so.

Monday, May 4, 2009

Is Tenure a Problem in Science Departments?

Somewhat belatedly, I wanted to comment on the New York Times Op-Ed by Mark Taylor that addresses some of the flaws of the current academic system and proposes some solutions. The central problems that Taylor highlights in his article are that academic departments are too isolated from one another and from the world, and that there aren't enough academic positions for all of the people who are getting doctoral degrees these days. Taylor makes some very good points, but his article is strongly influenced by his experiences in a humanities department, and I am not sure how relevant his suggestions are for a science department.

Astronomy suffers from many of the same problems Taylor describes. There are far more graduate students and post-doctoral researchers than there are tenured professorial positions, and yet students are trained as if they were all to be professors. However, science students are often not paying their own tuition (it is paid by the grant of a supporting professor) and there are more options for science students outside of tenure-track positions because of the many sources of external funding that support scientific research. Unlike the humanities, there are a variety of scientific programming and research positions for graduates who are not seeking professorial positions. These positions are aligned with the education students receive through their doctoral research.

This isn't to say that there is not a major problem in scientific disciplines concerning the ratio of student positions to professional positions. Rather, it is that the problem may not be as closely tied to tenure and the longevity of tenured professors as in the humanities. The problem may be that professional positions available to graduating students are being occupied by the students themselves. Scientific research in the United States relies on a large pool of skilled labor. Currently, this labor is being bought at well under market price in the form of cheap graduate student researchers. If more of these positions were filled by full-time research professionals, we might have a healthier employment system for scientific academia.

The problem is that this raises the price of research in the United States and may result in a reduction in the total number of projects (and therefore, researchers) that can be supported. From the perspective of researchers, this may be a healthier state of affairs--to not be misled into spending 5-7 years underpaid as a graduate student only to find that the only way to continue to do what you've been trained to do is to continue to be underpaid. But unlike in the humanities, science graduate students usually have not accumulated debt beyond their undergraduate education and they have been supported (however cheaply) through this process. The solution, then, may simply be to ensure that prospective graduate students in science are well-informed about what employment prospects they should expect after they file their dissertation.