Showing posts with label order-of-magnitude. Show all posts
Showing posts with label order-of-magnitude. Show all posts

Saturday, August 20, 2011

7 Gpeople


While at a conference in Istanbul, I went to the İSTANBUL ARKEOLOJİ MÜZELERİ, which was an absolutely fascinating archeology museum. Istanbul has featured prominently in the growth of civilization, and I struggled to keep track of the many different civilizations and cultures that occupied the region at one time or another. I had to find a youtube video to help me sort it out.

There was a very nice exhibit on Troia (Troy), which apparently really existed; it's ruins were unearthed in a farm field not too far from here. It was settled, destroyed, and resettled in 9 different epochs before being ultimately abandoned. Even back then, it seems that anything you dug up had a 1000-year history. Buildings were built on top of buildings. The reconstruction of the different settlements was interesting.

It was really hard for me to get my head around how small cities were in comparison to now. Istanbul currently has 13 Mpeople living in its greater metropolitan area. In 3000 BC, that was the human population of the world. The largest cities in antiquity were ~250 kpeople. It made me wonder if all of our advancements in technology and culture in the last couple hundred years could be attributed strictly to 1) more people to do the work and 2) longer tails on the normal distribution of people with various abilities.

So just how many people were there as a function of time? The log-log plot above from wikipedia shows current best estimates. Apparently, some 70 kyears ago, possibly as a result of a major volcano eruption, the human population was reduced to something on the order of 1000 to 10,000 "breeding pairs." Since then, the population rapidly recovered to several million, where it remained stable until agriculture was developed. This is all in a nice video tracing genetic migration via mDNA.


Since then, there has been exponential growth (a line on log-log plots) with a transition to a slower growth coefficient at ~400 BC. Occasional Black Plagues aside, the human population has increased dramatically. I remember hearing once that half the people who ever lived are alive right now. That's actually definitely false--it's closer to 6%. Also, everyone seems to think that population growth is accelerating (remember this video from the 80's ?). The graph above definitely shows that's not true, either.

But what is true is that this growth cannot continue unchecked without hitting its head on something, be it food supply, global warming, danger of pandemics, warfare, declining birth rates, or whathaveyou. It's estimated that in October of this year, 2011, there will be 7 Gpeople on the planet. This map shows where the population currently is, but it's estimated that much of the growth in the next century will happen in poverty-stricken Africa. Things pretty much have to plateau around 10 Gpeople, though.

So what to do? The most effective ways to reduce birth rates, which is key to controlling population growth, global warming, saving the environment, and many of the rest of our problems are:
  1. contraception
  2. improving the standard of living (ending poverty)
  3. education (and education about contraception)
  4. and reducing infant/child mortality
That last item is counter-intuitive. The reason it is important is that when survival rates are low, couples have more children to compensate, including a buffer for uncertainty. Having a predictable path from birth to adulthood allows for more precision in family planning.

Tuesday, December 28, 2010

Distribution of Actual Births Around Estimated Due Date

Given that we're now in overtime here for our second child, I have become very interested in the question of what our "due date" meant in the first place.

The plotted on the left are data from a study of Canadian births, 1972-1986 (Arbuckle & Sherman 1989). I then started trying to aggregate some data for my own plots.

The first thing I noticed required a little attention was the "fence post" problem relating to recording gestational age. If a study A records births in the 39th week, where exactly is that on the x axis? Well, if we assume that they started counting with 1, then the 39th week is actually 38.5 +/- 0.5 weeks (they are counting fence). But if study B records births from week 39-40, the implication is that they started at 0 (they are counting fence posts), so 39-40 is 39.5 +/- 0.5. That was a little tricky.

Then we have studies that bin over different time intervals. If another study C records births at 37-41 weeks, how do we relate that to A and B? The intelligent way to plot this would be to use the probability density of going into labor, which divides out by the length of time over which the observation was made. So we should be careful to do that.

And then there's just the general problem that a lot of studies list percentages for births in each time bin, but don't list total populations or error bars, so we don't know what the errors in their measurements were. Ugh. So I crossed my fingers and hoped they followed good practices with their significant figures: I assigned an error equal to the last significant digit they list.


I got most of my data from this semi-thorough compilation of census data and going to some of the original sources. The data aren't great, but they are adequate (see plot to left). I fit to the data an increasing exponential tail to the left, plus a normal distribution. The fit isn't great (despite some claims in the literature of it being normally distributed), but it captures enough of the overall distribution.


The width of the normal distribution was 1.5 weeks, centered at 39.5 weeks. This seems consistent with several sources that suggest ~10% of pregnancies would go into the 42nd week if they were allowed to. It also suggests that the "due date" means the mean of the normal portion of the distribution. Yet fully 1/2 of pregnancies will go beyond the expected due date, and 1/6th will go past 41 weeks, according to this coarse fit.

Of course, none of this accounts for biases that we know exist. The growing prevalence of inductions and C-sections move births earlier artificially. Although the statistical significance may questionable owing to systematic biases (self-selection for uncomplicated pregnancies, etc.), it appears that the recorded midwife births go later than the aggregate (presumably hospital-dominated) births at the 2-sigma level. This may potentially indicate that without intervention biases, the distribution of birth dates around the expected due date could be broader and weighted toward later dates.

Monday, March 29, 2010

Who Dominates Health Care Costs?

There's nothing like a little bout with MRSA to make one pay a little more attention to the state of health care legislation. Two nights in the ER are definitely making me thankful for health insurance. Knowing that it wasn't going to cost me an arm and a leg to get antibiotics through an IV (in fact, it probably saved me the leg) definitely helped me to seek care early, rather than waiting for the infection to get truly life-threatening. And that probably saved in health care costs in the long run.

I've heard it argued many times by the other side that universal health care will drive up the cost of health care for everyone, because so-called "healthy people" will be paying, through their premiums, for the bills of the "unhealthy". Ignoring that:
  1. the above is a tautological statement about what insurance is
  2. people routinely go from the "healthy" group to the "unhealthy" group and back again
  3. we should maybe feel a moral obligation to care for the unhealthy
Yeah, ignoring that, I wanted to know if the underlying assumption was, in fact, true. Who dominates health care costs? Is it the small number of extremely sick people? Or is it the larger number of moderately sick people?

To answer that question, I went searching for the population distribution of health care costs. I found the following publication: Variations in Lifetime Healthcare Costs across a Population (Forget et al. 2008). To the left are reproduced Figs. 4 and 5.

Given all the hype, I was somewhat underwhelmed to see that these curves depict (with the exception of an excess at the lowest cost bin) a gamma distribution. This isn't surprising, because a gamma distribution is supposed to represent the sum of a bunch of exponentially-distributed random variables.




To find the contribution of people in each cost bin to the total health care cost of the population, we simply need to multiply the population of that bin (drawn from a gamma function) by the mean health care cost of that bin (a linearly increasing function). Setting the mode of the gamma distributon to $90k for females, and tweaking the k and theta parameters (I'll chi-by-eye it at k=4.5, theta=1.0) we get the following distributions of fractional population (black) and fractional total health care cost (red), as a function of lifetime healthcare cost:

So who dominates health care costs? Those just slightly above the mode, which is to say, the large number of people who are just a little sicker than most. And that really could be any of us, folks.

Friday, January 8, 2010

Hands-On Cosmology Education

Yesterday I spent the morning giving a gosh-wow talk about cosmology to a physics class at Athenian High School taught by my housemate Dave Otten. It was a lot of fun, and the students were all very enthusiastic. It was almost entirely driven by their questions, and they loved being pitched curveballs (time is reference-frame dependent, the universe is expanding, spiral arms are standing waves, etc). The hour-and-a-half lecture was over before we knew it.

Afterward, Dave mentioned that it would be really cool if there were a way to talk about galactic-scale astronomy and cosmology that was in keeping with the philosophy of their school, which emphasizes lab-based, hands-on learning. He mentioned that PhET is a free resource he uses for providing interactive simulations that make hands-on labs out of subjects that otherwise would be too slow, small, big, fast, or dangerous to perform live in a classroom. He also lamented that there aren't any galactic- or cosmological-scale simulators there that could help to understand how systems on this scale behave, and that could perhaps illustrate exactly where the problems of dark matter and dark energy are encountered. Has anyone seen something like this?

Thursday, July 2, 2009

The Need for Speed


On the drive from San Juan to Arecibo this morning, I got to wondering about where my average driving speed fell in the distribution of drivers here in Puerto Rico. In the states, I felt like I was a pretty average driver, but here en la isla, the distribution of driving speeds is different. There are a lot of fast drivers, too be sure, but there is also a subpopulation of drivers whose speed is significantly (~10 mph) below the speed limit. This may be because relative to the US, PR is economically depressed and so more old cars are on the road, or as a reaction to the more erratic driving habits there seem to be here, but anyway, I definitely pass more people than pass me now.

So in an effort to discover where my driving speed fell relative to others (and in and effort to alleviate the boredom of driving 1.5 hrs alone), I started counting how many cars I passed and how many passed me as I was going 65 mph (the speed limit). Out of 55 pass events, only 9 involved me getting passed. To make this a tractable problem in my head, I decided to assume that driving speeds were normally (gaussian) distributed about a mean--even though this contradicts my anecdotal evidence above. Using this approximation, my first instinct was to say that 1/6 of the cars on the road were faster than me, and since ~2/3 of samples are within +/- 1 sigma of the mean in a gaussian distribution, 1/6 of the samples would be above +1 sigma. So I approximated that I was a 1 sigma driver.

But then it occurred to me that I needed to control for a significant sample bias. This is because the test I was doing wasn't randomly selecting cars and comparing my speed to them. Cars were far more likely to get selected if the difference between their speed and my speed was large. A car going the same speed as me would never pass me, and I would never pass it. But I would assuredly pass almost every car on the road that was going 10 mph as I went 65. The "road distance" that I sampled for different velocities is proportional to abs(v-v0), where v0 is my velocity. The effect this had on my samples was to underweight speeds close to my own and overweight the wings of the gaussian distribution I had assumed as my model. If I drove exactly the mean velocity, this effect would not be terribly important--if the model were correct, it would still be the case that as many cars passed me as I passed. But as my velocity moves away from the mean velocity, the "normal drivers" who are only going a little faster than me get undersampled, so I only see the drivers who are tearing around like a bat out of hell. At the lower end, I still see the real slow-pokes on the road, but I start seeing people who are going a bit faster than that, of which there are a lot more. The effect of this sample bias, it seems, would be to make it seem that I'm a farther outlier in my driving speed than I actually am.

So now, to figure out where I fall in the (normal) distribution of driving speeds, I need to know exactly what the mean driving speed and what sigma is, so that I can compensate for the abs(v-v0)
sampling factor. That means I need to figure out 2 numbers, but unfortunately, I only measured 1 number (that 1/6 of the passes while driving were me being passed) so I won't be able to properly constrain this problem. However, I should be able to figure out the mean on my drive home by finding the speed at which as many people pass me as I pass. For now, let's say this is 55 mph. Then all I need to do is find the sigma for which a gaussian distribution around 55 mph downweighted by abs(v-65 mph) has 1/6 of the area lying above 65 mph. I just solved that numerically on my computer, and it's saying that the best-fit sigma is ~22 mph. So that puts me at about +1/2 sigma. That seems reasonable.

An interesting next step (which I'm not going to do right now since I need to get to work) would be to translate the sample error in my pass measurements into an error in the determination of sigma, and then the error in my driving speed percentile.

Saturday, January 19, 2008

Rain Water

I consider myself an environmentalist. I try to do the basic things around the house that help reduce my carbon footprint (you can calculate yours here) like installing compact florescent lights, installing power strips to prevent appliances like TVs from pulling current while "off", driving the minimum possible, eating vegetarian, and drinking soy milk instead of the dairy variety (cows are surprisingly bad for the environment). And I'm always on the look-out for something crazy and interesting to try for the environment. An idea I've been tossing around in my head (especially now that I live in rainy Puerto Rico) is harvesting power from rainwater. Part of my recent interest in rainwater, I must admit, has to do with the fact that we average about one water-outage per month here; I want backup water. However, my original interest, a year or two ago, was actually in generating electricity from the rain falling on my roof.

Here are some order-of-magnitude calculations for what I might expect to get from this. Quick research indicated that average annual rainfall here is 62 inches (about 150 cm). This means that each square meter of rooftop collects about:

150 (cm) x 100 (cm) x 100 (cm) x 1 (g/cm^3) / 1000 (g/kg) = 1500 kg

of water per year. If a story of a building is 5 m high, then the potential energy in that water is:

Force x Distance = Mass x Gravity x Distance = 1500 (kg) x 10 (m/s^2) x 5 (m) = 75,000 (J) = .02 (kWhr)

So PR generates .02 (kWhr/m^2/story) annually. My apartment building is about 10 (m) x 20 (m), and is 4 stories high, bring the energy of rainfall to 16 kWhr annually. If electricity costs $0.10 per kWhr, I could save a whopping $1.60 off my power bill annually. That probably won't ever offset what it would take to build the generator. Sigh.

Of course, dams do exactly the above, but with an enormous collecting area (not just a rooftop--an entire drainage basin). It looks like rain on my roof isn't going to power my computer, though. But I still might try to use the water to flush my toilet when the water goes off next weekend.