Let’s have a task: in whatever way suits you best, note down the Facebook posts of those pages you follow whose content at least remotely resembles a reaction to, or a critique of, political events (in other words, no kittens and no sophisticated analyses of sports matches). Why?
Well, firstly, you will most likely be surprised that the selection of what you have liked over the past year is not as varied as you thought; yes, you are making it easier for various marketing tools to get a pretty clear picture of your personality with the help of clever algorithms. That machine learning and
Big Data
are already the present of the future and big business is no great news; some banks are even starting to go through your social data this way in order to refine their algorithms for calculating the risk of whether or not to give you that loan for a new Škoda. Lev Manovich, a superstar in the academic circles of new media, in turn pointed out on his Twitter that a technology company focused on Big Data can identify the consumer habits of individual users from the locations they visit and slot them into pre-prepared boxes of consumer types. For readers versed in the overcrowded terminology of the slightly amorphous field of User Experience Design, this is in fact the automated creation of
personas
, by which they even fulfil one of the basic conditions, as defined by the interaction designer Alan Cooper, that personas be based on real data from real users. On my Facebook the machine would come across such oddities as the fact that I calmly follow both Ádvojka and Deník Referendum and, for example, Parlamentní listy, indeed even Pravý prostor.
For technical reasons, however, I left the Czech media aside, since for the linguistic analysis I only had a (large) corpus of the English language available. After browsing through FB pages I selected these more or less extravagant sources of information:
- BBC News
- Bloomberg Business
- CNN Politics
- Counter Current News
- Democracy Now
- Deutsche Welle
- Haaretz (an Israeli newspaper)
- Mondoweiss
- The New York Times Opinion
- PRESS TV (an Iranian source in English)
- Reuters
- Salon
- The Economist
- The Guardian
- The Intercept
- The New Republic
- The Washington Times
- Truthdig
- Wall Street Journal
Thanks to Facebook’s Graph API I was able to download over 1,000 posts that the news outlets I follow shared on their walls over one week. I wanted to know which keywords appear most often in these news items on a given day and at a given hour. For the linguistic analysis of all these posts I used Python and the TextBlob library. For each post I went through the headline and the lead and recorded all the noun phrases, which I made into the keywords of that post.
The source data in JSON after the noun phrase analysis for each post:
source_data_json.7z
I then worked with the resulting JSON file in JavaScript, where I parsed the data and visualised it in a suitable way. I thought that for analysing the occurrence of keywords over a period of time some kind of slider would serve well, according to which I would display the data for the visualisation.
Whereas during the week the keywords were spread evenly across many topics, from the beginning of Friday night the reaction to the terrorist attacks in Paris starts to dominate very clearly.
You can judge the embryonic result for yourselves: