Friday, 31 January 2014

arXiv trawl: January 2014 - Social Media



An interesting way of keeping your ear to the ground regarding the latest happenings in the scientific world is to monitor the arXiv. The arXiv (pronounced "archive”) is a repository for electronic preprints of scientific papers. Scholarly peer review of scientific papers can take a long time, so many scientists use archives like this to share their findings and to seek comment on their work before official publication. As such, the content of the arXiv is many and varied; there are weird and wonderful topics, and papers in various states of review. Some will never get published anywhere else, whilst others are seminal (for example, Perelman’s proof of the Poincare conjecture). But by the very definition of preprint, they are all calling for comment. I dived in recently, and here are some highlights of the last few months on the arXiv concerning social media.

Because MySpace --> Facebook

Researchers from Princeton, in the report Epidemiological modeling of online social network dynamics, modelled the rise and fall of MySpace by likening it to a disease and used epidemiological methods to model how it infected the population, and how the population eventually became immune. The number of times the term “MySpace” was searched for in Google was used as a measure of the site's popularity (or how infected the population was). This data was sourced from Google Trends. If you would like to read more about the maths involved, check out Sick of Facebook? Read on… in Plus. As you can see below, they fitted a nice curve to the data. Cute.



All good so far. Stories concerning social media are favourites of conventional media, and naturally this was picked up: Facebook could fade out like a disease. What the newspapers focused on was the work fitting the epidemiological model to Facebook data (Google searches for “Facebook”) and the conclusion that Facebook is heading for a "rapid decline", and between 2015 and 2017 will lose 80% of the users.

There are two questions that arise from this:

1) Is Google Trends data a good measure of the popularity of a website?
2) Just because the MySpace data fits this curve does not mean Facebook will.

Facebook was made aware of this study, and their reply was pretty excellent. They did their own study using Google Trends on searches for "Princeton" and found that:

“Princeton will have only half its current enrollment by 2018, and by 2021 it will have no students at all, agreeing with the previous graph of scholarly scholarliness. Based on our robust scientific analysis, future generations will only be able to imagine this now-rubble institution that once walked this earth.

While we are concerned for Princeton University, we are even more concerned about the fate of the planet — Google Trends for "air" have also been declining steadily, and our projections show that by the year 2060 there will be no air left.”



Thanks to FlowingData for the link to Facebook's reply.

It’s also a nice example of a study that will never be followed up by the conventional media. Even if this work makes an entirely accurate prediction of Facebook’s future, there will be no follow up newspaper article in 2020.

How to get yourself retweeted, sort of

Everybody likes to be popular. To this end, Ronald Hochreiter and Christoph Waldhauser authored A Genetic Algorithm to Optimize a Tweet for Retweetability in which they look at the factors that make a tweet popular. They developed a Twitter-like network of connected people and pushed tweets into the network to see which ones were retweeted. They modified the tweets each time they ran the simulation using a genetic algorithm. These algorithms work like evolution. Each tweet had a number of “chromosomes” that were modified with random mutations each time the model was run, with those mutations that brought about better results (that is, more retweets) kept, and those that didn’t, further randomly mutated. The chromosomes of the tweet concerned the polarity of the tweet (I think this means whether it’s positive or negative, it’s not well explained), how emotional the tweet is, the length of the tweet, the time of day it’s sent, and the number of URLs and hastags contained.

So how should you compose your next tweet? Well, unfortunately the paper doesn’t really say. It shows some results that don’t translate particularly well into reality. Their conclusions are that the genetic algorithm works pretty well, and that more work is needed. Fair enough, there are plenty of papers out there that simply outline how a new model works rather than having exciting results; I’ve written a few myself.

Don’t share your mobile phone number on social networks

Call Me MayBe: Understanding Nature and Risks of Sharing Mobile Numbers on Online Social Networks wins this month’s award for best reference to a pop song in a scientific paper title. The researchers examined how sensitive personal information spreads around social networks by collecting 76347 unique mobile numbers posted by 85905 users on Twitter and Facebook. They then used these mobile numbers to gain sensitive information about their owners from other social networks.

This in itself is an interesting study of how easy it is to collect personal information online, but they didn’t leave it there. They then communicated the observed risks to the owners by calling them up with the mobile numbers they found. Some users were surprised to know about the online presence of their number, while others had intentionally posted it online for business purposes. They found that 38.3% of users who were unaware of the online presence of their number had posted their number themselves on the social network.

Where’s my hoverboard?

Searching the Internet for evidence of time travellers makes the bold claim in its abstract that, regarding the search for time travellers online, it is “perhaps the most comprehensive to date”. Essentially what they did was search the web for information that shouldn’t have been known at the time of publishing – only time travellers from the future could have possessed such prescient knowledge. The two events they were looking for evidence of were the viewing of Comet ISON and the inauguration of Pope Francis – both big events that people in the future would know and care about. To do this, they needed to look for information published before these events occurred. They found that Bing and Facebook were no good for this study as they didn’t make clear at what date the information was published, or the date could be easily edited. So they used Twitter, on which tweets are nicely time-stamped. They called for time travellers to use the hashtag #ICanChangeThePast2 in September 2013 and looked at tweets before this time. They also examined Google Trends for searches a time traveller might have made.

Disappointingly, they found no evidence that any time travellers concerned themselves with posting on twitter or doing google searches.

I am going to go for a run before I press publish on this post. So, if there are any time travellers out there, come and join me at Erskineville Oval at 1pm Thursday 30th January.

(Edit: There were two people and a dog down at the oval. The dog chased and barked at me in a very knowing fashion. The time-travellers of the future are apparently long haired, short brown dachshunds.)

References:
  • John Cannarella, & Joshua A. Spechler (2014). Epidemiological modeling of online social network dynamics. arXiv: 1401.4208v1
  • Ronald Hochreiter, & Christoph Waldhauser (2014). A Genetic Algorithm to Optimize a Tweet for Retweetability Proceedings of MENDEL 2013: 13-18. 2013. arXiv: 1401.4857v1
  • Prachi Jain, & Ponnurangam Kumaraguru (2013). Call Me MayBe: Understanding Nature and Risks of Sharing Mobile Numbers on Online Social Networks. arXiv: 1312.3441v1  
  • Robert J. Nemiroff, & Teresa Wilson (2013). Searching the Internet for evidence of time travelers. arXiv: 1312.7128v1

Saturday, 5 May 2012

Ep 145: Teleportation


Is teleportation possible in the real world, or only in the world of science fiction?

In this very special episode, Dr Boob takes the reigns and leads us on a journey through teleportation, whether or not physics allows it and even if it does, can we technologically achieve it? What are the implications if we recreate someone in another spot - what about their soul? Does such a thing exist? And even if you can technologically achieve this, is it possible to reanimate a copy of someone? What do you do with their original version, if you have simply copied them? This could be considered cloning, which brings in ethical questions.

Perhaps wormholes could be a solution to this problem, but we haven't found any yet - however they are, as physicists like to say, theoretically possible.

Tune in to this very entertaining episode (and I can say this without any false modesty as Dr Boob did it all himself) here.


If you'd like to hear more of Dr Boob on this podcast, check out our past joint episodes, mostly on the science of superheroes. He's also on twitter, so come and follow him, he needs friends!


Friday, 22 April 2011

Ep 141: Science of Superheroes - Harry Potter


And we're back! It's been a while, but finally it's time for another podcast, so we've made it a long one. Take this episode on a long train ride or car trip, as Dr Boob and I explore the science of the spells of Harry Potter.

Attempting to find scientific and engineering solutions to Harry Potter spells is probably the most difficult task we have set ourselves yet, so we would be very interested to hear how you would made the Harry Potter spells a reality. The spells dealt with in this episode are:
  1. Lumos - Producing light from the end of a wand (A voice activated torch seems a logical solution),
  2. Aguamenti - Shooting water from the end of the wand,
  3. Alohomora - Picking a lock at a distance,
  4. Expecto Patronum - Protection against evil dementors in the form of some virtual creature,
  5. Sectumsempra - Slicing your opponent open,
  6. Aparecium - Reading invisible ink,
  7. Accio - Summoning things to you,
  8. Expelliarmus - Disarming your opposition of their wand,
  9. Confundo - Confusing the victim,
  10. Stupefy - Stunning the victim,
  11. Invisibility cloak - Covering yourself in a cloak to make yourself invisible,
  12. Imperio - Forcing your victims to obey your commands,
  13. Obliviate - Erasing the memories of the victim,
  14. Legilimens - Telepathy.
Although some of these are quite clearly impossible at the moment, in every case we have come up with a scientific or engineering solution to take us at least part of the way there. Listen in to find out what we came up with, and please write in and let us know where we have gone wrong or what you would do.

Click play below or listen to this show here.



References:
  1. Santos, V., Paula, W., & Kalapothakis, E. (2009). Influence of the luminol chemiluminescence reaction on the confirmatory tests for the detection and characterization of bloodstains in forensic analysis Forensic Science International: Genetics Supplement Series, 2 (1), 196-197 DOI: 10.1016/j.fsigss.2009.09.008
  2. A.J. Barnier and D.A. Oakley (2009). Hypnosis and Suggestion Encyclopedia of Consciousness DOI: 10.1016/B978-012373873-8.00038-4
  3. T.C. Jerram (1982). Hypnotics and sedatives Side Effects of Drugs Annual DOI: 10.1016/S0378-6080(82)80009-3
  4. Wood, B. (2009). Metamaterials and invisibility Comptes Rendus Physique, 10 (5), 379-390 DOI: 10.1016/j.crhy.2009.01.002

Monday, 14 February 2011

Search Traffic in Egypt

The Egyptian Revolution of 2011 was a series of street demonstrations that demanded the overthrow of the Egyptian President Hosni Mubarak. One of the government retaliations to the protests was to shut down the Internet. Imagine you're a youth in Egypt and all of a sudden you don't have access to the Internet with its social networks, games and unlimited porn. You'd protest too! Great strategy guys...

Here is the Google search traffic in Egypt normalised against world-wide Internet traffic, created through Google Transparency Report. As you can see, it took a little less than a week for the government to realise their folly.

Wednesday, 1 December 2010

2D / 3D / 4D Baby Ultrasounds

Being able to see your unborn child is truly an amazing experience. Ultrasound (diagnostic sonography) is a common diagnostic tool for, among other things, imaging the foetus to determine its age, look for abnormalities and observe blood flow in the umbilical cord. But possibly its most memorable effect is seeing your baby's heart beat - and in 3D/4D ultrasounds, seeing your baby's face.

The term "ultrasound" applies to acoustic energy (sound) with a frequency above the audible range of human hearing (20 Hz -20 kHz). When used in medical imaging, an ultrasonic sensor (or transducer) is placed on the mother's belly and produces pulses of sound. The frequencies used for medical imaging are generally in the range of 1 to 18 MHz. High frequencies (7-18 MHz) can be used to look for fine details but have low penetration, so to image deep tissue, lower frequencies (1-6 MHz) are used.

The sound waves are partially reflected from layers between different tissues inside the mother's body. Sound is reflected anywhere there are density changes - for example, at the baby's skin where it meets the amniotic fluid. The baby's internal organs can also be imaged depending on what frequencies you use. The reflected sound is then "heard" by the transducer, and the data analysed to produce the image. The amount of time it takes for the echo to rebound relates to how deep the sound penetrated, and the strength of the return signal relates to both the material it is reflecting off and its depth. The deeper the tissue from which the signal is being echoed, the quieter the return, simply because there is more sound loss (attenuation) the further the sound travels (it gets absorbed, scattered and reflected along the way). This information allows an image to be built up, whereby pixels at the appropriate depth are coloured by the strength of the return at that point. Generally, the sound waves are not 100% reflected at any stage - you can see "behind" objects because some sound penetrates through. However, as less sound is penetrating the deeper you go, the signals become fainter.

2D Ultrasounds

Baby face 2D scan

The typical ultrasound image is a "2D" image like the one above. In this image, the transducer is at the top and is sending sound waves down. The image is essentially a slice through the mother. It's called a 2D image as we can only see two dimensions - left/right and up/down. The 2D image is built by firing a sound beam down, waiting for the return echoes, and then firing a new pulse at a slightly different angle. This continues until an arc is swept. Combining the data from each line after the arc is swept gives the 2D image. The following images come from the excellent resource Basic ultrasound, echocardiography and Doppler for clinicians, by Asbjorn Støylen. The left image shows the transducer scanning whilst the right image shows how the pulses are sent down in lines.



Continual rescanning means that a 2D video can be produced with roughly 50 frames per second. The human eye can see about 25 frames per second and so the video looks smooth. This frame rate is also more than enough for 2D temporal visualisation of the baby's heartbeat (~70-150 beats per minute depending on age) and to watch blood flow through Doppler ultrasound. Due to the Doppler effect, the sound pulse will rebound with a higher frequency if it hits something moving towards it, and a lower frequency if it echoes from something moving away from it - this is the same reason the noise of a car has a high pitch when moving towards you, and a low pitch as it moves away. As blood is moving in the umbilical cord, the ultrasound can be coloured by the Doppler information to show the blood flow.

3D Ultrasounds

Baby face

3D images are a fairly recent advance in diagnostic sonography. Instead of just seeing a slice through the mother, the images can show a surface - essentially adding depth (the third dimension) to the 2D image. Imagine you are looking at a car from front on - you have no idea how long the car is and you have no information on how many doors it has or if the boot is open. However, if you look at the car from another angle, you can figure this out, and the more angles you look down, the more depth information you can gain. This is essentially what a 3D ultrasound does - it stitches together multiple 2D shots from different angles to produce the image. Modern transducers have the ability to scan multiple cross-sections. If the baby is moving, there may be some blur, but as image processing is becoming quicker, the 3D images are becoming clearer. The colour of the image is not real as there is no way to see colour inside the mother. 3D scans provide information for the diagnosis of facial anomalies, evaluation of neural tube defects, and skeletal malformations, and also helps the parents bond with their unborn child (it's very cool). However, when compared to 2D scans, they aren't as useful for the diagnosis of congenital heart disease and central nervous system anomalies. One of the reasons why this is the case is because they are static, which leads us to...

4D Ultrasounds

The term 4D refers to the addition of time to 3D scans. This is a very recent advance as it is only in the last few years that we have had the computing power to not only stitch together the 2D images to make the 3D images, but to create the 3D images quickly enough to play them consecutively as a video. Modern 4D scans play at roughly 12 frames per second, so they are a little jumpy.

Here is a little video I put together of our 4D scan.



I don't know if there is an upper bound on what ultrasound technology can do - as the speed of sound is ~1540 m/s in human soft tissue, and you have no choice but to wait for the return signal before you can process the image, it may be that a high video frame rate with decent resolution is unobtainable. Resolution depends on how many different lines you fire down to make the first 2D image - more lines mean better resolution, but currently you have to wait for the echo from one line before sending down the next, which means it takes longer to produce an image. I imagine one way of improving this would be to send down all the lines at once with slightly different frequencies or waveforms, and as such when the echo is received you would know where it came from. Perhaps this is already being done - let me know if you know more!

Check out the video of Massive Attack's Teardrop in which there is a singing foetus, and I also have more images over at my ultrasound set on flickr.

References:
  • Kurjak, A., Miskovic, B., Andonotopo, W., Stanojevic, M., Azumendi, G., & Vrcic, H. (2007). How useful is 3D and 4D ultrasound in perinatal medicine? Journal of Perinatal Medicine, 35 (1), 10-27 DOI: 10.1515/JPM.2007.002 
  • Carrera, J.M. (2006). Donald School Atlas of Clin. Application of Ultrasound in Obs/ Gyn www.jaypeebrothers.com DOI: 10.5005/jp/books/10226
  • Khanem, N. (2007). Donald School Textbook of Ultrasound in Obstetrics & Gynecology The Obstetrician & Gynaecologist, 9 (2), 140-140 DOI: 10.1576/toag.9.2.140.27325 

Wednesday, 21 July 2010

Ep 132: Science of Superheroes - The Hulk

The science of superheroes is taking a green and nasty turn this week as we discuss the largest superhero of them all, The Hulk. Join myself and our regular superhero expert Dr Boob as we delve into the science of how we might realise The Hulk in the lab. It was one of the more entertaining interviews I have done for the podcast.

Listen in to this show here (or press play below), and read further for more info:



The Hulk is alter-ego of Dr Bruce Banner, who allegedly bares a striking resemblance to Dr Boob. Banner is a reserved physicist who involuntarily transforms into The Hulk when triggered by a strong emotion such as anger, fear, terror or grief. The Hulk himself is a massive green monster who gets stronger the angrier he gets. He also has bullet-proof skin.

The Hulk’s origin story includes depends on whether we are looking at the comic book Hulk, the Hulk of the two recent movies, or The Incredible Hulk of the TV series (in which it is David Banner, not Bruce Banner, who metamorphoses into The Hulk).

The 2003 movie version "Hulk" includes many of the topics we discuss in the podcast. The movie starts with genetics researcher David Banner – Bruce Banner’s father - working with the military to "improve" human DNA. The opening credit sequence depicts experiments with jellyfish and starfish DNA, and Banner’s notepad mentions bioluminescence. This suggests that the Hulk gets his green colour from jellyfish DNA as some jellyfish bioluminesce at around 450 nm, which is at the blue/green end of the spectrum. In 1961, Osamu Shimomura extracted green fluorescent protein and another bioluminescent protein, called aequorin, from Aequorea victoria while studying bioluminescence. He eventually received the Nobel prize in Chemistry in 2008 for this work. The mention of starfish is also interesting because, as we found with Wolverine, starfish and sea cucumbers have great healing powers and are able to regenerate lost limbs. Evidently, Banner wanted to splice bioluminescence and improved healing into human DNA.

Banner’s experiments then moved to lizards and monkeys, but unfortunately they all died. Naturally, he then decided if his experiments did not work on animals, he would try them on himself – clearly, ethics committees are not part of superhero science. After conducting experiments on his own DNA, he eventually passes on his mutant DNA to his unborn son Bruce. Once David realises this, he changes his approach and works to cure his son of his genetic afflictions, however the research is shut down and an explosion kills David’s wife. David is taken to a lunatic asylum and Bruce is adopted.

Years later, Bruce has followed his father’s line of work and is conducting military research – Bruce’s area of interest is the use of nanomeds in soldiers. This might include such things as targeted drug delivery for rapid recovery from injury. An experimental accident subjects Bruce to an enormous dose of gamma radiation which “activates” his mutant DNA (possibly combining with the nanomeds) and the building rage/stress transforms him into The Hulk for the first time.

Whether or not this is scientifically possible – well, that’s the topic of the podcast so tune in!

Other issues that we discuss include:
  • Gamma radiation and radiation poisoning;
  • Genetic transfer and gene therapy – could David Banner change his own DNA in such a way that this change would be copied to his progeny? For more information, check out the Weismann Barrier;
  • The Hulk’s size – is it possible to rapidly increase your size? Simple conservation of mass equations would suggest no, and bacteria in a Petri dish generally have a 24 hour doubling time. There are also enormous metabolic requirements involved – we need to have resources available to feed these growing cells and Bruce Banner is not excessively fat. Perhaps to do this we need to accelerate Bruce Banner to the near the speed of light, at which point he may relativistically pick up some mass - however, this is not particularly practical!
  • The Hulk’s strength – is it possible to rapidly increase your strength?
  • The Hulk's healing properties - could we use some of the science of Wolverine here?
  • The materials used to create bullet-proof skin. The toughest skins in the animal kingdom are crocodile, elephant, shark and armadillo; however none are bullet (and knife) proof;
  • What materials could we use to make his "one-size-fits-all" pants? You will notice that no matter what size Bruce Banner or The Hulk are, and no matter what the ripped state of his other clothes, his undies always fit.
  • And of course, whether The Hulk has irritable bowel syndrome and wears giant green snuggies.
    Hope you enjoy this show - we certainly enjoyed recording it, as you will be able to tell by the end! Listen in to this show here (or press play below):



    NB: I've now discovered there's a Red Hulk - future show perhaps?
    Samples in this podcast are broadcast courtesy of ioda PROMONET. They were:

    The Toxic Avenger
    "Superheros 2007" 
    from "Superheroes" 
    Buy at iTunes
    Spaceman
    "Superhero"
    from "Little Baby Souls"
    Buy at iTunes
    Candye Kane
    "Superhero" 
    from "Superhero"  
    Buy at iTunes
    Ninja Kodou
    "Superhero (Psychedelic Man)"
    from "Ninjutsu"  
    Buy at iTunes

    References:
    Shimomura, O., Johnson, F., & Saiga, Y. (1962). Extraction, Purification and Properties of Aequorin, a Bioluminescent Protein from the Luminous Hydromedusan,Aequorea Journal of Cellular and Comparative Physiology, 59 (3), 223-239 DOI: 10.1002/jcp.1030590302

    Moghimi, S. (2005). Nanomedicine: current status and future prospects The FASEB Journal, 19 (3), 311-330 DOI: 10.1096/fj.04-2747rev

    Friday, 18 June 2010

    How do you spell goal?

    It's not often I get the chance to pursue three of my passions - sport, mathematics and online social media - at the same time. The 2010 Football World Cup combines these facets of life in a way we probably haven't seen before, providing numerous opportunities for data mining, funky visualisations and general nerd-indulgence, as well as knocking out twitter for a time.

    One creative exploration I saw recently was by @neilkod, who collected data from 30 GB of tweets on how people spelt the word "goal". Data mining Twitter is an evolving field - see our recent story on how by using Twitter data you can predict the success of a film. You can find the full goal data table here, and I have listed the top and bottom few below:

    Rank
    Word
    Count
    1
    goal
    50225
    2
    Goal
    11727
    3
    GOAL
    4202
    4
    goAl
    798
    5
    goall
    340
    6
    GOAAAL
    92
    7
    goaL
    88
    8
    Goall
    75
    9
    GoaL
    69
    10
    GOAl
    66
    11
    goaaaal
    61
    12
    GOal
    50
    ....
    ....
    ....
    1249
    GGGGGGGGGGGGGGGGOOOOO
    OOOOOOOOOOOOOOOOOOOA
    AAAAAAAAAAAALLLLLLLLLLLLLLLL
    1
    1250
    GGGGGGGGGGGGGGGGGGGOO
    OOOOOOOOOOOOOOOOAAAA
    AAAAAAAAAALLLLLLLLLLLLLLLLL
    1
    1251
    GGGGGGGGGGGGGGGGGGGOO
    OOOOOOOOOOOOOALLLLLLLL
    1
    1252
    GGGGGGGGGGGGGGGGGGGGG
    GGGGGGGGGOOOOOOOOOOO
    OOOOOOOOOOOOOOOOOOOO
    AAAAAAAAAAAAAAAAAAAAAAA
    AAALLLLLLLLLLLLLLLLLLLLLLLLL
    1

    As expected, on top is the word "goal" (71%) followed by "Goal" (17%) and "GOAL" (6%). Then there are various misspellings, before the excited tweets come in, including the 140 character "Goooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooal".

    Interestingly, the caps lock induced "gOAL" is only used once.

    You can also visualise this in a chart - note that in the chart below, the x-axis is a log scale due to the fact that the leading terms are so far ahead:

    Distribution of the word "goal"

    Now it's time to get our nerd on....

    Zipf's law is a curious law that arose out of an analysis of language by linguist George Kingsley Zipf, who theorised that given a large body of language (that is, a long book), the frequency of each word is close to inversely proportional to its rank in the frequency table. That is:


    where a is close to 1. This is known as a "power law" and suggests that the most frequent word will occur approximately twice as often as the second most frequent word, which occurs twice as often as the fourth most frequent word, etc. There is an excellent take on this over at Plus Magazine. As you can see in the log/log chart below, after the 5th version of "goal" a Zipf curve fits remarkably well.

    Distribution of the word "goal" - log/log chart

    There has never been a real explanation of why Zipf's law should apply to languages and there is controversy surrounding whether it gives any meaningful insight. Power laws relating rank to frequency have been demonstrated to occur naturally in many places - the size of cities, the number of hits on websites, the magnitude of earthquakes and the diameters of moon craters have all been shown to follow power laws.

    Wentian Li demonstrated in his paper Random Texts Exhibit Zipf's-Law-Like Word Frequency Distribution, published in IEEE Transactions on Information Theory, that words generated by simply randomly combining letters fit the Zipf distribution. Li showed mathematically that the power law distribution of frequency against rank is a natural consequence of the word length distribution, with words of length 1 occurring more frequently than words of length 2 and so forth. His underlying theory is that the rank distribution arises naturally out of the fact that word length plays a part - long words tend not to be very common, whilst shorter words are. Li argues that as Zipf distributions arise in randomly-generated texts with no linguistic structure, the law may be a statistical artifact rather than a meaningful linguistic property.

    Our results mirror Li's quite closely. It is clear that the most used versions of the word "goal" - where goal is spelt correctly with various capitalisations - should not fit the Zipf distribution as these words are not random - they are the actual correct spellings people are looking to write in their tweets. However, after the initial few words, random spelling errors, and then the simple randomness of how long people hold their fingers on the keys in the excitement of a goal, take hold. From this point, we see exactly the same as Li - that the Zipf distribution arises from random words, with longer words less common than shorter words.

    If you would like to see some truly insightful twitter world cup visualisations, check out The Guardian's World Cup 2010 Twitter replay. I haven't been replaying the Australia vs. Germany game very often....

    Wednesday, 26 May 2010

    Ep 130: Using Twitter to predict Movie Box Office Revenue (and the future...)

    Ever wondered what good Twitter actually does? Personally, I love it, but really, is it anything but noise?

    One of the pipe dreams for online social media is the ability to track opinions and interests in real time. In their paper Predicting the Future With Social Media, Sitaram Asur and Bernardo A. Huberman have not only tracked live opinion on movies, but used it to predict their future success.

    Asur and Huberman, from the Social Computing Lab, HP Labs California, have shown that the rate of tweeting about a movie accurately predicts its opening weekend box office revene.

    After examining the rate of chatter from almost 3 million movie tweets, the researchers constructed a linear regression model for predicting box-office revenues of movies in advance of their release. These results outperformed the Hollywood Stock Exchange, a market in which people can buy and sell virtual shares in actors, directors and individual movies and produces unusually accurate predictions of film popularity.

    There is a strong correlation between the amount of tweets concerning a forthcoming film, and its opening weekend box office return. The rate of tweeting about a movie was determined by simply counting the number of tweets containing the movie name. The next step was to predict box office returns beyond the opening weekend, and this was achieved by including "sentiment" as a factor. Sentiment analysis is a fascinating area of linguistic study. Language classifiers were used to label the text associated with the movie tweet as Positive, Negative or Neutral. Adding these as factors into the regression significantly increased the researchers' ability to predict the box office returns beyond the opening weekend.

    These results are intuitive - before a movie is released, potential viewers do not know whether they will like the movie and so simply the number of tweets about a movie gives an indicator of movie "buzz" and correlates with the number of people attending the opening weekend. Once a movie is released and people start forming opinions, movie tweets start to contain sentiment. Negative tweets, whilst they have little effect on the opening box office as no one has yet seen the film, have a strong influence on further returns. Likewise for positive tweets.

    The question of cause and effect is very interesting. Does a high number of tweets about a movie actually cause a strong box office return, or are they correlated simply because the twitter and movie audience are arguably the same? Another way of asking this question is to ask whether an advertiser could change future box office returns by deliberately tweeting multiple times or with a particular sentiment.

    I had a fascinating chat with Sitaram about this work. To listen to our chat, tune in here (or press play below):



    References:
    Sitaram Asur, & Bernardo A. Huberman (2010). Predicting the Future with Social Media arXiv.org arXiv: 1003.5699v1

    Wednesday, 19 May 2010

    I want a Turbo Encabulator!

    I also want some Prefamulated Amulite

    Monday, 12 April 2010

    Ep 126: Science of Superheroes - Doc Ock

    Continuing with our recurring segment The Science of Superheroes, this week we're tackling the mechanically-blessed supervillain Doc Ock, from Marvel Comics. And joining me once again for a journey through superhero scholarship is Dr Boob. To listen to this show, tune in here (or press play below):



    Dr. Otto Gunther Octavius is a scientist who designed a set of advanced mechanical arms to assist him with his nuclear physics research. He controlled the arms via a brain-computer interface. In the movie Spiderman 2, Octavius created the mechanical arms to help him conduct nuclear fusion experiments. The arms had their own artificial intelligence, with an inhibitor chip used such that Octavius could maintain control over them.

    The arms attached to a harness that was strapped around his body. In great comic book tradition, a freak experimental accident caused the limbs to fuse to his body, and the inhibitor chip was destroyed. The arms themselves took control as Octavius could no longer control them, and mad-scientist Octavius became evil Doc Ock. Interestingly, the limbs were able to defend themselves whilst Doc Ock was unconscious, implying not only self-awareness, but a capability to sense their surroundings.

    In this episode, we come closer than we have come before in our series to figuring out a way to recreate a superhero (or supervillain in this case) in the laboratory. The topics discussed in this podcast include:
    • Robotics,
    • The history of artificial limbs,
    • The history of aritifical intelligence, and how to design limbs that could possibly have self-awareness and a desire (and capability) to defend themselves,
    • What is nuclear fusion? Is it possible to develop a controlled energy source using nuclear fusion, and if so, could this be the way forward for powering enormous artificial limbs?
    • What would the limbs be made from? Is it time to turn once again to Adamantium? See our show on Wolverine for more information.
    • Assuming the AI is difficult to accomplish, how could the limbs be controlled? Two methods include:
    1. Myoelectric prostheses - a myoelectric prosthesis uses EMG signals from muscles on the surface of the skin to control the movements of an attached prosthesis. These prostheses have been used where arms and legs have been amputated, with the prosthesis attaching to the residual limb. The concept of neuroplasticity is also very important here. Neuro- (or brain-) plasticity is the ability of the brain to change throughout life, to reorganise itself and form new connections between neurons. Artificial limbs have recently been controlled by chest muscles - this is an example of the brain learning how to control muscles in a completely new way.
    2. Remote control - recent work has shown that objects can be remotely controlled by brain waves (EEG). Naturally, this does not mean one can levitate a chair on the opposite side of the room - the brain needs to be hooked up to a computer which reads the brain signals, interprets them and then controls the connected object in an appropriate way. We discussed this a few years ago in our article Space Invaders Mind Control, Small Testes and Facial Expressions

    To listen to this show, tune in here (or press play below):



    And on the topic of superheroes, you may enjoy this poster from Russell Walks Illustration. It is a Periodic Table of 122 fictitious elements from sci-fi movies, comics, TV series etc. Adamantium is in there, as naturally is the most famous of them all, Kryptonite. Click on the image below for a closer look and to buy the image as a poster.



    Thanks to @markfromhouston for the tip on the Periodic Table poster.

    Saturday, 13 February 2010

    Ep 122: Science of Superheroes - Wolverine (Part 2)

    This is the second part of our series on the science of Wolverine - specifically, how can we create Wolverine in the lab? Join Dr Boob and myself as we journey through Wolverine's characteristics and how they may be recreated in a human. Read more on Wolverine in part 1 of this series. To listen to this show, tune in here (or press play below):



    Specifically in this episode, we tackle the topics of:
    1. What would happen to your bones if you completely covered them with metal? Bones are living parts of your body and make red blood cells, platelets and bone marrow - among other things - that are vital for life.
    2. Would a lack of platelets reduce Wolverine's ability to heal?
    3. Wolverine is likely to be on a cocktail of drugs, including anabolic steroids to beef him up, immunosuppressants so his body doesn't reject the metal coating on his bones, and various drugs to supply red blood cells, bone marrow and platelets.
    4. Could we really harness the healing powers of the sea cucumber for Wolverine, and would they work quickly enough?
    5. Are carrots enough to improve his sight?
    6. What metal could we use to coat his bones? It needs to be able to be injected as a liquid and then harden at body temperature. Most steels have melting points over 1000 degrees Celcius, and this would cause terrible trauma to his body. Dr Boob's suggestion was CerroLOW117, which is 44.7% Bismuth, 22.6% Lead, 8.3% Tin, 5.3% Cadmium and 19.1% Indium. CerroLOW117 has a melting point of 47 degrees Celcius, however lead and cadmium both accumulate in the body and have adverse health effects. It is highly likely CerroLOW117 would not be strong enough to help Wolverine anyway.
    7. And what is a phlebotomist?
    For more on superheroes, check out our recurring science of superheroes series. And for more from Dr Boob, check out Chris's other contributions.

    Monday, 25 January 2010

    Keyboard cat tops the charts

    Internet memes are fascinating. The term meme refers to ideas and cultural phenomena that spread through society through imitation. The term was first used by Richard Dawkins in his book The Selfish Gene to discuss how evolution could work with cultural phenomena such as beliefs, fashion and music. He argues that memes are cultural analogues to genes, in that they self-replicate, mutate, can be inherited and respond to selective pressure.

    The concept of Internet memes relates to this original definition of meme, and refers to the spread of ideas across the internet. Colloquially, internet memes are essentially in-jokes. And like any good in-joke, the meme will often start out as not particularly funny, but after you have seen it 1000 times, it takes on a life of its own, often evolving to something completely unpredictable. Lolcats is an example of a meme that I don't think is all that funny - however, I am quite partial to Rick Rolling, and I Like Turtles is awesome!

    I was recently putting together my annual list of most played songs on my ipod for the last year when I realised that I am clearly susceptible to internet memes. Not only was the list Rick Rolled (with no less than three Rick Astley songs) but the top song was one I learnt of through my favourite internet meme, Keyboard Cat. Keyboard Cat consists of 1984 footage of "Fatso" the cat playing the keyboard, and was uploaded to youtube as charlie schmidt's "cool cat". The gag is to append footage of the cat to video of people doing stupid things (falling over, getting hit in the head etc.) - the cat is essentially playing the person off stage, much like getting the hook in the days of vaudeville.

    One of the best Keyboard Cat videos - and probably my favourite video ever on youtube - has Keyboard Cat with Hall and Oates playing their song You Make My Dreams. The video starts with a segment of Desperate Lives, a 1982 movie starring Helen Hunt showing the effects of drug use - Keyboard Cat plays off the overdosing Hunt as a cautionary tale against drug use. Then the music video starts, and this is when I discovered the song - check it out!



    2009 saw a minor comeback for the song, and even though I have no evidence to back this up, I am going to put it out there that Keyboard Cat is the reason for its renewed success. Here's You Make My Dreams from the movie (500) Days of Summer - video here.



    And for a complete understanding of Keyboard Cat, check out the excellent Know your Meme series of videos.

    Monday, 11 January 2010

    Nano Snowman

    In 2009 we had Nano-teddy, this year it's Nano Snowman!

    David Cox from the National Physical Laboratory in the UK has built the world's smallest snowman - perhaps because he was trapped inside by the winter blizzards with nothing much else to do...

    The snowman is 10 µm across, which is roughly 1/5th the width of a human hair. Its head and body are made from two tin beads. The eyes and smile were carved from the tin with a focused ion beam, and the nose, which is less than 1 µm across, is a deposited blob of platinum.

    Hope all our UK and European readers are enjoying the snow!

    Thursday, 30 July 2009

    Nano Teddy!


    1stplaceHeliaJalili
    Originally uploaded by rodolfobernal
    I love this pic from flickr. It is a scanning electron microscope image of Zinc Oxide (ZnO) nanostructures (that is, really small ZnO structures) on indium oxide coated glass.

    I'm not sure if it is a chance photo, or one deliberately created, but either way it is cool!

    Scanning electron microscopes create images by scanning a surface with a beam of high-energy electrons. The electrons interact with the atoms that make up the sample, and this interaction produces various signals from which a sample's surface topography can be recreated. The interactions between the electron beam and the sample can give off x-rays or light, it can also cause the electron beam to be scattered, and small currents can be generated in the sample. These signals can be read by sensitive equipment which can then reconstruct what the surface looks like.

    See this "Science as Art" competition for more funky science photos!

    Wednesday, 29 July 2009

    Distracted Driving and Cut-throat Capitalism

    Here are a few games that are fun and sciencey to keep you sane at work.

    Driving whilst Distracted


    The New York Times recently published a Distracted Driver game to gauge your distraction while you're texting on the road. The game puts you driving on a road having to negotiate a number of toll booths along the way. The game tests your ability to drive through the correct gates without any distractions, and then it makes you write a couple of text messages whilst still having to negotiate the booths.

    After you finish the game, you get a comparison of your result with everyone else who has played. I improved the second time I played, my first results are below. I didn't see a grey lady either time!

    The science of mobile phone use whilst driving is a developing field, with most of the research suggesting that you are just as impaired, or more so, if texting or using a hand-held mobile as you are if you are drunk. A couple of great resources if you are interested are:
    1. The Dutch national road safety research institute (SWOV). Their publication SWOV Fact sheet: Use of mobile phone while driving was published in 2008 and contains a great deal of up-to-date research. Their conclusion is that the negative effects of mobile phone use whilst driving are caused by both physical and cognitive distraction. Although physical distraction can be reduced through the use of such aids as handsfree phones and speed dialling, cognitive distraction remains the crucial problem. They conclude that handsfree phones do not have significant safety advantages over handheld phones. They also point towards research suggesting that talking on a mobile phone is associated with cognitive distraction that may undermine pedestrian safety.
    2. Applied Cognition Laboratory, Department of Psychology, University of Utah - David Strayer has published a wealth of research on the impact of using in-car technologies on driving performance and traffic safety. It is well worth a browse of their published articles.

    Cut-throat Capitalism


    Piracy has a romantic history often associated with walking-the-plank, peg-legs and saying Arrrrr a lot - there is even an international talk like a pirate day on the 19th of September each year. However, modern piracy in Somalia is a deadly game and nothing like the stereotype. Nevertheless, Wired Magazine has had some fun with this and brings us Cut-throat Capitalism in which you are a pirate commander staked with $50,000 from local tribal leaders and other investors, and your job is to guide your pirate crew through raids in and around the Gulf of Aden, attack and capture a ship, and successfully negotiate a ransom.

    The game is addictive and highly strategic. Initially I kept alienating my crew by being nice in my negotiations, so they eventually deposed me as captain. Then when I was too tough on my hostages, the Navy was called in. Eventually, I was successful in negotiating a $3 million ransom.

    Piracy off the Somalian coast is seen as a business by those who conduct it, and as such it can be analysed by economists (who, in general, love to think that the world revolves on an economic axis). Wired has taken a look at the economics of piracy and found that the typical payoff for piracy in Somalia today is 100 times what it was in 2005. One of the reasons why it is flourishing today is because it exploits the incentives that drive international maritime trade. Shippers, insurance companies, private security contractors, and national navies stand to lose less by tolerating it rather than attempting to stop it - insurance covers the ransom, ships and goods are not lost and no one dies. The pirates know that they can keep escalating the situation to see just how much the "market" can bear. The negotiation process also involves risk/reward calculations.