Saturday, 31 May 2014

Some life analysis with Twitter

There was a great post recently on Flowing Data, The Change My Son Brought, Seen Through Personal Data. It got me thinking about what my life looks like through personal data, and probably the best source of data since the advent of smartphones is Twitter. Twitter recently made it possible to download your personal archive and it makes for some interesting analysis. Along with RSS feeds, Twitter is my major source of online news, education and entertainment, and it is also useful for personal communications and microblogging.

Downloading your personal archive is easy, but you need to do a little manipulation before you can play with it. My tweets were time-stamped in UTC time (I'm not sure why - perhaps by default, perhaps because of my location settings) so I had to adjust this for time zone changes due to day-light savings and overseas trips (I didn't bother with domestic trips as I don't have an easy record of them, and they don't make too much difference - an hour here and there).

The following has a dot for every tweet I've written since the end of 2010. Take note that the x-axis is quite long (3.5 years) and the dots are quite large (bigger than a day). I haven't annotated it, but it is interesting to spot life events - the birth of my children, various periods of leave and holidays, over-tweeting during The Ashes etc. There are auto-tweets that came out at the same time each week (which I've now stopped as they're annoying). There was a definite shift in the time I rise in the morning after December 2010 when my son was born and a surge in late-night tweets after my daughter was born in 2013.


Breaking it down is a little more interesting. The following shows tweet frequency for work days and non-work days (weekends, leave) since the start of 2013. On a work day, I tweet in the main on the train. I usually catch a train around 7 or 8am in the morning and the return train around 5 or 6pm. During work hours there is a trickle through coffee breaks and lunch, and after dinner is another peak. This type of profile aligns somewhat with the findings of other social media studies (Yellow Social Media Report – 2014 - thanks @problogger), although the amount I tweet on the train is more than the norm, whilst the amount I tweet at work is less (although it is a great way to horizon scan the various fields of science in which I work, once you follow the right people).

Non-work days follow a different profile, at least until after dinner. There's a slightly later rise in the morning, dips when we would be attempting to get out of the house, a dip at an earlier dinner time and a large peak in the evening once the kids are in bed. This peak is higher than a work day, in which time I might be preparing for the next day or falling asleep on the couch. By about 10pm it is basically the same till 6am the next day.


I'm posting this at about 9am on a weekend, having written it at about 10pm last night - that fits the curve pretty well. If you are a social media marketer (of which, at last count, there are 1,083,645,638 on Twitter), target my work trips (although I'm sure you know this from all that stunning big data analysis you do). The downside of this is that the train trip is too short to read anything of any length, which would explain why the 140 characters of Twitter spikes at these times.

Saturday, 21 July 2012

Visualising Runs

Inspired by a recent post from Kasey Clark in which he plotted all his runkeeper runs (tracked via GPS) on a single map, I thought I'd explore my own running from the last few years and see how it might be visualised in an interesting manner.

Using his method, I exported all my runs as one big zip file of gpx files (found under your profile) then imported them all into Google Earth. Here is an image of all my runs around Sydney's inner west over the last few years. Most of the time I run along the Cooks River.



I also had a bit more fun with it, and for this you will need the Google Earth plugin for your browser - if you can see the following images you already have it, and if not then there should be a link for you to get it.

The city2surf is one of the world's biggest fun runs and I have done it the last few years. By creating a Google Earth tour, you can create an animation of your runs. I tweeked the gpx code in a text editor (and Excel) to make my 2010 and 2011 runs start at the same time, and then by using the tour gadget, you can embed the animation on your website. Perhaps over time I will add further year's runs to this animation. You'll need somewhere to host the exported kml files from Google Earth. There is a small lag at the start of the video and if it doesn't work, see the video on youtube. I'm looking to knock off that 2011 time this year in a few weeks! Edit 1: I have added a friend from 2010 and 2011.
The next tour doesn't look so great but it would look great in San Francisco or New York City. Google Earth has 3D buildings built in, and by turning these on, you can visualise your runs in 3D. The following shows my Bridge Runs across the Sydney Harbour Bridge and finishing at the Opera House. Runkeeper doesn't quite get the elevation of the bridge correct so it looks like I'm running across water. As mentioned, in cities where there are lots of rendered 3D buildings, this would look great. I haven't bothered yet to tweek the start times for each of the races to all be exactly the same as it's a bit fiddly, but you get the point. Again there is a small lag and if it doesn't work, see the video on youtube.
If you can't see the above videos, and the Google gadget seems really buggy, I have uploaded them to youtube and there you can see city2surf and bridge runs videos.

Sunday, 12 February 2012

The Big Swim

Recently I competed in one of Australia's biggest ocean swims, The Big Swim. Now I'm not particularly good, just stupid and competitive, and the results provide a nice sporting dataset with which to play. I've wanted to teach myself some mapping / visualisation techniques for a while, so I took the opportunity to investigate this data in order to find out from where competitors for the event came, and from where they are the quickest.

I have created the following interactive chart using Google Fusion Tables. From the swim results, I extracted the competitors' times and the suburbs they came from, and then mapped the suburbs to their postcode using the aus-emaps postcode finder. From this table I worked out the average, minimum, maximum and median times for each postcode. I've only plotted New South Wales postcodes.

The tricky part was mapping the postcode boundaries. Thankfully, the Australian Bureau of Statistics has a couple of files you can use, however to use these with Google Maps, you need to convert them to the kml file type. MyGeodata Converter provide such a service. This meant we had two files - one with the swimmer statistics per postcode, and one with the boundary coordinates. It is easy to merge these tables with Google Data Fusion, and voila, you have an intensity map.

The map below is coloured by the number of competitors from each postcode - red is the most and green the least. The most swimmers came from postcode 2026, which is Bondi and surrounds. Many postcodes, including my own, only had one competitor. If you click on a postcode, it will give you that postcode's statistics - note that the times are in decimal (Google Data Fusion has some issues with data type, so it was easiest to treat the times as decimals, rather than date/time format). So 51.58 minutes means 51 minutes 35 seconds.

The quickest postcode (that had over 10 competitors) was 2075 (St. Ives and surrounds). The slowest with over 10 competitors was 2153 (Baulkham Hills and surrounds). One might postulate that Baulkham Hills is too far from the beach, and that everyone in St. Ives has a private swimming coach. Or it could just be random, as there really aren't enough swimmers per postcode to draw too many conclusions.

The biggest bug in this is the "Sydney" postcode which is, I'm fairly sure, way over populated due to people putting "Sydney" down instead of their suburb in their swim registration. Not that many people live in the city.



The following chart shows the distribution of times, which looks quite like a normal distribution with a slight right skew due to the fact that there is a hard limit on the quickest you can possibly complete the swim, whilst you can take as long as you like to finish. Large public sporting events tend to have a long tail as people may come out once a year and jump in the ocean without particularly caring how quickly they go. This is especially true for running events where you often have people dressed up as Snoopy out the back. Ocean swim events tend to have less of this as, unlike running, if you stop, you drown! So without a very long tail, the Central Limit Theorem kicks in and gives you a normal-ish (or log-normal distribution) distribution.


References:
  1. The results come from the Ocean Swims website (which is an excellent source of information for ocean swimming in Australia) - the Ocean Swim Series website is also a good data source.
  2. Make your own tables and maps at Google Fusion Tables.
  3. The postcode information came from the Australian Bureau of Statistics and aus-emaps.
  4. I converted the ABS data to a kml file using MyGeodata Converter.
  5. All Things Spatial is a great resource for data mapping

Monday, 14 February 2011

Search Traffic in Egypt

The Egyptian Revolution of 2011 was a series of street demonstrations that demanded the overthrow of the Egyptian President Hosni Mubarak. One of the government retaliations to the protests was to shut down the Internet. Imagine you're a youth in Egypt and all of a sudden you don't have access to the Internet with its social networks, games and unlimited porn. You'd protest too! Great strategy guys...

Here is the Google search traffic in Egypt normalised against world-wide Internet traffic, created through Google Transparency Report. As you can see, it took a little less than a week for the government to realise their folly.

Friday, 18 June 2010

How do you spell goal?

It's not often I get the chance to pursue three of my passions - sport, mathematics and online social media - at the same time. The 2010 Football World Cup combines these facets of life in a way we probably haven't seen before, providing numerous opportunities for data mining, funky visualisations and general nerd-indulgence, as well as knocking out twitter for a time.

One creative exploration I saw recently was by @neilkod, who collected data from 30 GB of tweets on how people spelt the word "goal". Data mining Twitter is an evolving field - see our recent story on how by using Twitter data you can predict the success of a film. You can find the full goal data table here, and I have listed the top and bottom few below:

Rank
Word
Count
1
goal
50225
2
Goal
11727
3
GOAL
4202
4
goAl
798
5
goall
340
6
GOAAAL
92
7
goaL
88
8
Goall
75
9
GoaL
69
10
GOAl
66
11
goaaaal
61
12
GOal
50
....
....
....
1249
GGGGGGGGGGGGGGGGOOOOO
OOOOOOOOOOOOOOOOOOOA
AAAAAAAAAAAALLLLLLLLLLLLLLLL
1
1250
GGGGGGGGGGGGGGGGGGGOO
OOOOOOOOOOOOOOOOAAAA
AAAAAAAAAALLLLLLLLLLLLLLLLL
1
1251
GGGGGGGGGGGGGGGGGGGOO
OOOOOOOOOOOOOALLLLLLLL
1
1252
GGGGGGGGGGGGGGGGGGGGG
GGGGGGGGGOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOO
AAAAAAAAAAAAAAAAAAAAAAA
AAALLLLLLLLLLLLLLLLLLLLLLLLL
1

As expected, on top is the word "goal" (71%) followed by "Goal" (17%) and "GOAL" (6%). Then there are various misspellings, before the excited tweets come in, including the 140 character "Goooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooooal".

Interestingly, the caps lock induced "gOAL" is only used once.

You can also visualise this in a chart - note that in the chart below, the x-axis is a log scale due to the fact that the leading terms are so far ahead:

Distribution of the word "goal"

Now it's time to get our nerd on....

Zipf's law is a curious law that arose out of an analysis of language by linguist George Kingsley Zipf, who theorised that given a large body of language (that is, a long book), the frequency of each word is close to inversely proportional to its rank in the frequency table. That is:


where a is close to 1. This is known as a "power law" and suggests that the most frequent word will occur approximately twice as often as the second most frequent word, which occurs twice as often as the fourth most frequent word, etc. There is an excellent take on this over at Plus Magazine. As you can see in the log/log chart below, after the 5th version of "goal" a Zipf curve fits remarkably well.

Distribution of the word "goal" - log/log chart

There has never been a real explanation of why Zipf's law should apply to languages and there is controversy surrounding whether it gives any meaningful insight. Power laws relating rank to frequency have been demonstrated to occur naturally in many places - the size of cities, the number of hits on websites, the magnitude of earthquakes and the diameters of moon craters have all been shown to follow power laws.

Wentian Li demonstrated in his paper Random Texts Exhibit Zipf's-Law-Like Word Frequency Distribution, published in IEEE Transactions on Information Theory, that words generated by simply randomly combining letters fit the Zipf distribution. Li showed mathematically that the power law distribution of frequency against rank is a natural consequence of the word length distribution, with words of length 1 occurring more frequently than words of length 2 and so forth. His underlying theory is that the rank distribution arises naturally out of the fact that word length plays a part - long words tend not to be very common, whilst shorter words are. Li argues that as Zipf distributions arise in randomly-generated texts with no linguistic structure, the law may be a statistical artifact rather than a meaningful linguistic property.

Our results mirror Li's quite closely. It is clear that the most used versions of the word "goal" - where goal is spelt correctly with various capitalisations - should not fit the Zipf distribution as these words are not random - they are the actual correct spellings people are looking to write in their tweets. However, after the initial few words, random spelling errors, and then the simple randomness of how long people hold their fingers on the keys in the excitement of a goal, take hold. From this point, we see exactly the same as Li - that the Zipf distribution arises from random words, with longer words less common than shorter words.

If you would like to see some truly insightful twitter world cup visualisations, check out The Guardian's World Cup 2010 Twitter replay. I haven't been replaying the Australia vs. Germany game very often....

Tuesday, 8 June 2010

Visualising the Music Universe

Every now and then I like to post data visualisations, and this one comes from one of my favourite web applications, Last.fm. Last.fm remembers what songs you play on your iPod (or whatever music player you use). Given that we now have an electronic device ban in our workplace, I'm giving their online streaming radio a go. (Ask me what I think about this ban and I'll give you a forthright answer, offline...) In any case, last.fm is a data lovers dream. The following visualisation shows my listening habits in 2009. The planets represent the top tags given to songs I listened to throughout the year - generally, these tags are song genres. Planet size is proportional to how much of that tag I listened to. The moons represent my top artists of 2009, with the light side showing how much I listened to the artist this year and the dark side showing last year's listening. Moons that are clustered together suggest that those artists are similarly tagged. The orbits show the correlation between the artist and the tag - artists with a closer orbit are more strongly associated with that tag than artists on the outer orbits.



Very funky, even if sometimes the moon is larger than the planet, and the small planet "Old People" appears in my solar system - you'll need to check out the high resolution picture to make out the smaller planets and moons.

Wednesday, 10 February 2010

Tracking Mouse Clicks

Here is another funky app to add some interest to your work day. Anatoliy Zenkov has developed a tool (Mac and PC) to track your mouse pointer during the day. Simply open it up and let it run in the background for as long as you like. 8 hours of mouse tracks from my day today look like this:

My mouse tracks


The big dots are times when the mouse did not move - for example, the bottom right dot is my lunch break. I've added some extra notes to the image on flickr to explain these dots. My movements around the taskbar are clear, as are mouse movements to the top left - I was using Excel for most of the day, and many of the functions I used were located there, as were the open / close file options. I found this tool on FlowingData. I'd be interested to see your tracks.

Tuesday, 17 March 2009

Correlation of the Week: Intelligence and Music Preference

The other day I heard loud distorted music approaching from a hotted-up 1986 Holden Calais with mag-wheels and a ridiculously loud sub-woofer and thought:
  1. I wonder what this guy is compensating for, and;
  2. His music taste probably says a lot about his intelligence.
Well, the study has been done. Before we jump into it, it's worth saying that the author of this study, Virgil Griffith from California Tech, makes no claims about correlation equalling causation. He just presents his results and allows us draw our own conclusions. His method of correlating music with intelligence involved:
  1. Find the ten most frequent "favourite music" descriptions at every US college via that college's Network Statistics page on Facebook;
  2. Download the average SAT/ACT score from CollegeBoard for students attending those colleges;
  3. Correlate the Facebook music results with SAT/ACT results and draw your own conclusions on music taste and intelligence.
The artist associated with the highest intelligence, by a clear margin, is Beethoven, whilst some rapper by the name of Lil Wayne seems to be loved by those less blessed in their mental faculties. The study also showed that Counting Crows, Sufjan Stevens, Radiohead and Ben Folds Five appealed to big brains whilst very disappointingly, for me anyway, Beyonce is at the other end of the scale. According to the data, people who listen to "indie" music are the smartest, and the genres come out:

Soca < Gospel < Jazz < Hip Hop < Pop < Oldies < Reggae < Alternative < Classical < R&B < Rap < Rock < Country < Classic Rock < Techno < Indie

It's tempting to think that intelligence has a direct impact on music choice, but this is probably not true - I know plenty of research scientists into Britney Spears. And the reverse - that music-choice influences your intelligence - doesn't make sense either, even though you may be occasionally tempted to think that listening to mindless dance-music makes you stupid. Could there be some drivers that influence both intelligence and music choice? Possibly. You can imagine that socio-economic factors and what you are exposed to whilst growing-up would influence the music you like and how well you do at school. Your parents would be a big influence too - I just can't shake Wet Wet Wet... There are countless possibilities that are best mulled over at the pub.

Whatever the case, I'm heartened by the results! We've already done a story on visualising music tastes using Last.fm, and most of my favourite artists are in the top half of the table with indie my favourite genre. For more on science and music taste, check out the podcast we put out in 2006 - one of the very first Mr Science Show episodes down the phone to China Radio International - called Can Scientists Predict Your Music Taste which looks at how web applications such as Last.fm and Pandora recommend songs to you based on your listening habits.

Griffith's study on music and intelligence comes on the heels of his "books and intelligence" study, in which he correlated book tastes with intelligence again using Facebook data. Harry Potter is the most popular book with The Bible second (for some reason, The Bible and The Holy Bible are different books). Some of the results include:
For more on the book study, check out booksthatmakeyoudumb. And for more on the music study, see musicthatmakesyoudumb. The following picture is one of the ways Griffith visualised his results. See where your favourite artists lie.

So congratulations to Virgil Griffith and your study on music tastes and intelligence, you have won Correlation of the Week - the Flying Spaghetti Monster would be proud!

Tuesday, 17 February 2009

The words of 2008

I recently stumbled across this wonderful visualisation tool called Wordle.

Using Wordle, I have created this image of the most popular words on The Mr Science Show blog throughout 2008 (not including common words like "the" and "and".)

The words most used on the Mr Science Show blog throughout 2008

It's nice to see that science is number one on the list! The image is quite a nice reflection of my interests in 2008, with maths and mathematical words such as distribution, stats and one featuring. We have sporting words such as cricket, league and sport, a few words artistic words such as music and dance, and some that need no explanation - sex and condoms....

I've started to become a little addicted to Wordle, so here is our DSTO Operations Research Code of Best Practice document - looks like a new funky ad campaign for studying mathematics!

The words most used on the Mr Science Show blog throughout 2008

Thursday, 9 October 2008

Last.fm, data mining and mashups

I've recently been putting together a Guide to Web 2.0 for The Helix Magazine and one of the most interesting aspects has been exploring the various mashups and applications of Last.fm.

Last.fm is a brilliant online music service and currently my favourite "web 2.0" application. By downloading a plugin for itunes (or whatever music player you have) that "scrobbles" each song you play (that is, tells Last.fm what you are listening to), a picture of your music taste builds up, and people with similar listening tastes are found. Artists are recommended to you according to your tastes, charts of your songs built up and "radio stations" perfectly tailored to you can be streamed online. But it is better than radio as there are no ads and you like every song.

By the way, I am westius on Last.fm.

Millions of songs are scrobbled every day by Last.fm users. This data helps Last.fm develop a massive database of user music preferences, and because of it's API, it is possible to access Last.fm information and develop interesting tools.

As users can tag their music with genres that they think aptly describe their songs and artists, it is possible to determine your own tag cloud of musical preferences. Using an excellent script at anthony.liekens.net, I came up with my own tag cloud, as you can see here.

It is possible from such tag clouds to examine how listeners fall into different categories through a process known as Data Mining. Data mining is essentially the process of sorting through enormous amounts of data and picking out the relevant stuff. Using principal components analysis - a mathematical technique which reduces multidimensional data sets to lower dimensions for analysis - and k-means clustering - an algorithm to cluster n objects into k groups - Liekens came up with 5 broad groups of Last.fm listeners:
  1. Electronic/pop
  2. Rock
  3. Indie
  4. Metal
  5. Hip-hop
Clearly this list does not reflect everyone on Last.fm (where are the classical music listeners?), but it does reflect the majority. I was surprised that Indie is a group in itself and am intrigued by the bundling of electronic and pop together - there are some tweaks to the maths you can make that could come up with different groups, and better results might be possible with a bigger data set . Hip-Hop listeners were the most clearly defined group. You can read more about the maths and how these groups are separated in the original article.

Another interesting thing you can do is compare your music tastes to your friends. This pic is a difference cloud comparing my music tastes with that of my good friend intranation. We have a roughly 40% similarity in music genre tastes, with the green tags those that I have more of in my collection, and the red those genres that intranation listens to more than me. No real surprises there.

Mashups are all the rage at the moment. The term refers to web applications that combine data from more than one source into a single integrated tool. For instance, domain, an Australian real-estate site, adds data from Google Maps to provide location information. My current favourite Last.fm mashup is idiomap. idiomap is a digital music magazine that personalises its content according to your interests in music, which it learns from your Last.fm profile. It gives you stories and reviews of the artists and genres you like, helps you discover new music and mashes in video and audio from youtube and other sources. idiomag aggregates music articles from over 100 different sources. You can also tweak the articles you like so if you receive something you don't like, you won't get it again. I subscribe to the RSS feed of my personalised idiomap magazine and so far its been great and has included reviews of music DVDs of artists I like and schedules of when bands will be playing and appearing on TV. Good stuff.

I will probably put out a few more blogs like this as I explore this world of mashups. And for podcast listeners, yes hopefully I will get one of them out soon too!