Showing posts with label babies. Show all posts
Showing posts with label babies. Show all posts

Saturday, August 20, 2011

7 Gpeople


While at a conference in Istanbul, I went to the İSTANBUL ARKEOLOJİ MÜZELERİ, which was an absolutely fascinating archeology museum. Istanbul has featured prominently in the growth of civilization, and I struggled to keep track of the many different civilizations and cultures that occupied the region at one time or another. I had to find a youtube video to help me sort it out.

There was a very nice exhibit on Troia (Troy), which apparently really existed; it's ruins were unearthed in a farm field not too far from here. It was settled, destroyed, and resettled in 9 different epochs before being ultimately abandoned. Even back then, it seems that anything you dug up had a 1000-year history. Buildings were built on top of buildings. The reconstruction of the different settlements was interesting.

It was really hard for me to get my head around how small cities were in comparison to now. Istanbul currently has 13 Mpeople living in its greater metropolitan area. In 3000 BC, that was the human population of the world. The largest cities in antiquity were ~250 kpeople. It made me wonder if all of our advancements in technology and culture in the last couple hundred years could be attributed strictly to 1) more people to do the work and 2) longer tails on the normal distribution of people with various abilities.

So just how many people were there as a function of time? The log-log plot above from wikipedia shows current best estimates. Apparently, some 70 kyears ago, possibly as a result of a major volcano eruption, the human population was reduced to something on the order of 1000 to 10,000 "breeding pairs." Since then, the population rapidly recovered to several million, where it remained stable until agriculture was developed. This is all in a nice video tracing genetic migration via mDNA.


Since then, there has been exponential growth (a line on log-log plots) with a transition to a slower growth coefficient at ~400 BC. Occasional Black Plagues aside, the human population has increased dramatically. I remember hearing once that half the people who ever lived are alive right now. That's actually definitely false--it's closer to 6%. Also, everyone seems to think that population growth is accelerating (remember this video from the 80's ?). The graph above definitely shows that's not true, either.

But what is true is that this growth cannot continue unchecked without hitting its head on something, be it food supply, global warming, danger of pandemics, warfare, declining birth rates, or whathaveyou. It's estimated that in October of this year, 2011, there will be 7 Gpeople on the planet. This map shows where the population currently is, but it's estimated that much of the growth in the next century will happen in poverty-stricken Africa. Things pretty much have to plateau around 10 Gpeople, though.

So what to do? The most effective ways to reduce birth rates, which is key to controlling population growth, global warming, saving the environment, and many of the rest of our problems are:
  1. contraception
  2. improving the standard of living (ending poverty)
  3. education (and education about contraception)
  4. and reducing infant/child mortality
That last item is counter-intuitive. The reason it is important is that when survival rates are low, couples have more children to compensate, including a buffer for uncertainty. Having a predictable path from birth to adulthood allows for more precision in family planning.

Tuesday, December 28, 2010

Distribution of Actual Births Around Estimated Due Date

Given that we're now in overtime here for our second child, I have become very interested in the question of what our "due date" meant in the first place.

The plotted on the left are data from a study of Canadian births, 1972-1986 (Arbuckle & Sherman 1989). I then started trying to aggregate some data for my own plots.

The first thing I noticed required a little attention was the "fence post" problem relating to recording gestational age. If a study A records births in the 39th week, where exactly is that on the x axis? Well, if we assume that they started counting with 1, then the 39th week is actually 38.5 +/- 0.5 weeks (they are counting fence). But if study B records births from week 39-40, the implication is that they started at 0 (they are counting fence posts), so 39-40 is 39.5 +/- 0.5. That was a little tricky.

Then we have studies that bin over different time intervals. If another study C records births at 37-41 weeks, how do we relate that to A and B? The intelligent way to plot this would be to use the probability density of going into labor, which divides out by the length of time over which the observation was made. So we should be careful to do that.

And then there's just the general problem that a lot of studies list percentages for births in each time bin, but don't list total populations or error bars, so we don't know what the errors in their measurements were. Ugh. So I crossed my fingers and hoped they followed good practices with their significant figures: I assigned an error equal to the last significant digit they list.


I got most of my data from this semi-thorough compilation of census data and going to some of the original sources. The data aren't great, but they are adequate (see plot to left). I fit to the data an increasing exponential tail to the left, plus a normal distribution. The fit isn't great (despite some claims in the literature of it being normally distributed), but it captures enough of the overall distribution.


The width of the normal distribution was 1.5 weeks, centered at 39.5 weeks. This seems consistent with several sources that suggest ~10% of pregnancies would go into the 42nd week if they were allowed to. It also suggests that the "due date" means the mean of the normal portion of the distribution. Yet fully 1/2 of pregnancies will go beyond the expected due date, and 1/6th will go past 41 weeks, according to this coarse fit.

Of course, none of this accounts for biases that we know exist. The growing prevalence of inductions and C-sections move births earlier artificially. Although the statistical significance may questionable owing to systematic biases (self-selection for uncomplicated pregnancies, etc.), it appears that the recorded midwife births go later than the aggregate (presumably hospital-dominated) births at the 2-sigma level. This may potentially indicate that without intervention biases, the distribution of birth dates around the expected due date could be broader and weighted toward later dates.

Tuesday, July 14, 2009

The Voynich Manuscript: Bootstrapping Language

The internet is a powerful and dangerous thing. It all started when I read the latest xkcd, which warned me that visiting cracked.com was dangerous. Then I read about the "6 phenomena that science can't explain" (which was a very dramatic title for some underwhelming mysteries), and next thing I knew, I was reading about the Voynich Manuscript. I learned about cryptography, glossolalia, the Manchu language, among other things. Then I took a look at the manuscript and before I knew it, I had a transcribed version of the manuscript in electronic form using the European Voynich Alphabet. And it just went downhill from there.

To summarize, the Voynich Manuscript (hereafter VMS) is a handwritten text with some illustrations some 500 years old. It uses glyphs no one knows how to read, it is not clear if it corresponds to a known language, it may or may not be encrypted, and little progress has been made in deciphering any of it, despite the fact that some bright people have tried. So I decided to have a crack at it.

The reason I got interested is because of the similaries to SETI. Arecibo, back in 1974, transmitted a message off into space that had been designed to be decrypted. We might receive a message like that some day. Or we might intercept something much like the VMS--a bunch of data in a language that we have no prior knowledge of--and we may be finding ourself trying to figure out how to bootstrap a language. That is to say, to learn the grammar and semantics of a language from a static example, without outside help.

Is this possible? For grammar, I'm pretty sure of it. I can imagine an algorithm (maybe Maximum Entropy Modeling and Bayesian learning applied to grouping and parsing) that uses correlations in the appearances of language elements (starting with letters and building up) and correlations in the behaviors of these elements relative to one another to build a model for parsing a language. For the VMS, I used something similar to this (not the MEM and Bayesian part) to show that spaces, line breaks, and paragraph breaks show similar grouping correlations relative to other VMS letters, and so can probably be considered one grammatical element of whitespace. That's a pretty simple thing to deduce, but it was actually something I was worried about in getting started with the VMS.

Sematics is another issue. Once upon a time, I would have had an optomistic answer to bootstrapping sematics from text that was not written for that purpose. However, after watching my child mysteriously acquire language, illustrating how hard-wired the human brain is for learning language from another human, and how much it relies on shared experience and feedback, I'm less sure.

I would be interested to know if there is a field of mathematics that studies sematics and the properties that a self-contained system needs to have to be able to generate sematical relationships. The Arecibo message relied on a shared physical environment to try to bootstrap sematics. I wonder if it would be enough to describe the rules of the grammar of a language in the language itself. That one, once the reader had deduced the relationships between elements, you would have a shared knowledge of that subject that might enable a reader to correlate the structure of the descriptions with the grammatical structure and thereby establish the first sematical relationships.

Anyway, after preliminary analysis of the VMS, I'm pretty sure that it's not random gibberish (there are correlations between elements on levels ranging from letters to words), and if it's encrypted, it's a weak form of encryption that preserves these correlations. My pet theory, extended from the glossalalia idea, is that this is actually plaintext in a natural language with an invented set of symbols, but that the natural language might be the accidental or intentional creation of a savant or scholar.

Saturday, October 4, 2008

Artificial Intelligence, Language Recognition, and Babies

As you may or may not know, I have son who was born 6 (almost 7) months ago. He is just the most incredible thing. Besides excelling in the normal metrics of cuteness and snugglability, he is also doing something that all babies do, but is probably the most incredible of all--he is is learning. He is learning how to move, how to recognize patterns, how to read expressions, how to form phonemes, how parse sounds. He is training the most incredible neural network on the face of the earth to solve problems that the best minds in the world have been working on for decades and haven't come close to solving. And he's making it look easy.

How is he doing it? Why can a baby, who knows nothing of the world, who lacks a fundamental understanding of physics, biology, optics, machine learning, linguistics, language, categorization, (the list goes on...) succeed at these tasks when very bright people using supercomputers cannot? I have some theories--some from an intro linguistics class I once took, some from personal experience trying to code speech recognition, some from learning foreign languages, some from my signal processing background, and some from watching this little ball of wonder over here. For what it's worth, here are my thoughts on the matter.

First, some observations:

  1. The problems listed above (motor control, image/speech recognition, etc.) are HARD. Just because the human brain is extremely adept at solving them, let's not make the mistake of underestimating their complexity.
  2. Babies don't come out with the answers. They pretty much can't do anything in the beginning. On the other hand...
  3. Babies have a definite propensity for arriving at a solution. They may not consciously know what they are doing, but they have a pre-programmed "boot sequence" that gets them walking, talking, and causing trouble by age 2. This boot sequence is remarkably consistent between babies (no baby walks before babbling, etc.)
  4. Babies have trainers (parents) who are instrumental in their development. However, babies work on problems of their choosing--a parent aids in language acquisition, for example, but cannot get a baby to start babbling before they come to it themselves. You can see this all the time when you watch kids. They have incredible attention spans for the skills they are working on, but things outside of that range are summarily ignored.
  5. Babies do not reason their way to solutions. Reasoning comes later.
  6. Changing gears... large neural networks don't work. Small neural networks are very good at discriminating between patterns on a few set of inputs, but one can't throw 1e4 pixels into a neural network and expect to train it to recognize any old picture of a cat.
  7. A lot of skills that we think of as a single skill (say, speech recognition) are actually many interrelated skills. For example, it is certainly my experience in learning foreign languages that: a) without adequate vocabulary, I have trouble hearing the sounds that are being spoken, b) without an understanding of what a conversation is about at a high level, I have trouble knowing what words to expect, and c) without being able to hear the sounds that are being spoken, I have no clue what a conversation is about. I'm not just being silly here. There's a real, circular dependence to speech recognition that requires several skills to be developed in parallel (phoneme recognition, vocabulary, grammar, cultural expectations) in order to advance.
  8. And finally, humans have an incredible knack for categorization--grouping things by common traits, and defining groups at all sorts of levels of generality.

I believe the above observations are only consist with the idea of a modular mind with a very strong hierarchy. In order to overcome the fact that large neural networks are untrainable, the brain has to be divided into modules that trained at particular sub-tasks that require fewer inputs. The mere fact that the brain is built of neurons and that these neurons only have several inputs suggests that this must be so.

Furthermore, the division of the brain into these modules must be pre-programmed. CAT scans reveal that the same physical locations in everyone's brain are responsible for certain functions, and it's clear that reason, cognition, and other general-purpose processing in our brain are not primarily responsible for language, image recognition, or motor control (although in adults, sometimes skills like language acquisition and motor control are augmented by reasoning and cognition). Humans have a natural language instinct, and capacity for image processing and motor control that belies an underlying, inherent cerebral architecture addressing these skills.

So what is the upshot of all of this? I think that work on artificial intelligence needs to reflect strong modularity and hierarchy. While designing hierarchical processing is not hard, training the kind of multi-tiered, cross-linked, sometimes circularly dependent system that AI requires is. Why do babies go through the same boot sequence? Do the stages of child development reflect the trained of different tiers of neural networks in the brain's hierarchy? It seems possible to me that the wonder of the human brain might more than this incredible hard-wired signal processing architecture--a fundamental component might be this incredible boot sequence, taking 10-20 years to complete, that trains neural networks ranging from simple movement and stimulus response to abstract thought and language.