Intro

For the course Foundations of New Media Studies we focused this time on social networks, which allowed us to explore and illuminate any topic we liked by means of data visualisation. As the tool for the analysis itself we chose the open-source visualisation program Gephi, which is built on the Java NetBeans platform. From it, Gephi inherits several good and bad properties: modularity, a large set of features, occasional instability and, last but not least, a GUI that was not designed as native to any specific OS, which may mean less than intuitive controls for many users. The authors of Gephi further boast that it is the Photoshop of data visualisation in its category. I leave it to the reader to judge whether that is a good slogan from a marketing point of view.

The topic was up to us, and so I decided to find out how people on the social network Twitter react to events that take up most of people’s free time, and not only on the other side of the Atlantic: the US presidential election.

Once I had chosen the topic, I faced a much more difficult question, one that was to decide whether the following hours of work would be merely aimless or useful playing at being a data analyst — the question of how I would handle the data, what data I would need at all, what I would be looking for in that data, and whether it is theoretically possible to get it out of that data. In other words, knowing what exactly to look for in the data is in many cases more important than knowing how to achieve something. In fact, these questions cannot be separated from each other, for they inform one another.

For example, excellent domain knowledge is, according to current theories of creativity, one of the preconditions for a person coming up with something new, useful and functional, which, incidentally, are the properties every creative output must have. Domain knowledge is, of course, a necessary precondition for eminent performance in any human activity. And it is precisely domain knowledge that makes data analytics much more than just the application of rigorous mathematical methods to a dataset; data analytics also requires the kind of knowledge the academic community considers soft: an awareness of the society, the culture and the current context of the area I am going to study, that is, a kind of humanities, soft-science awareness. My entirely speculative hypothesis explaining why there are so few good data analysts is precisely the uncommon requirement to be a good mathematician and a humanist at the same time, to have an education in computer science and mathematics combined with a dose of healthy awareness of what is going on in society outside the window, and then also a knowledge of sociology, anthropology, indeed even literature or philosophy, and, God forbid, new media too. It takes a poet and a chess player, a scientist and an artist.

All right, so how did I deal with this wicked problem, where there are many possible solutions, where the answer does not oscillate between a binary YES and NO but lies rather on a continuous scale of suitability? Which question did I decide to explore?

Twitter is a suitable platform for text analysis, above all because posts are consistent thanks to their 140-character limit, which simplifies the analysis. Besides the post itself, which I will call here the post proper, a Twitter post also contains text that serves as a description of the post proper; it is thus metadata. This includes the combination of the at sign @ + username of the user whom the post proper concerns in some way, and of course also the # sign + any text, a combination that is nowadays widely called a hashtag.

Because hashtags are often used not as an entirely irrelevant, unrelated string of characters but as metadata that supplement the post proper with useful contextual information helping readers understand the Twitter post, I decided to use precisely hashtags to arrive at a view of how people react to, write about and thus think about the presidential campaigns of the candidates I chose. But how exactly should hashtags be useful for this, and what does a hashtag actually mean?  

The social Web 2.0, folksonomy and hashtags

Etymologically and linguistically speaking, the word hashtag was formed by combining the words “hash” and “tag”. While the word hash refers uncomplicatedly to the hash or pound sign #, the second half of this compound, the word tag, is much more interesting, above all because tags were an important aspect of the transformation of the Web into a social platform, which came to be called, collectively, Web 2.0. The term was once used so often that many dismissed it as a mere marketing buzzword; nevertheless, it was and still is a useful label for the Web at the point when new technologies and new approaches to building websites and applications gave rise to tendencies and projects that differed noticeably from what, after version 2.0 was introduced, we call Web 1.0. A summary of all the significant changes was given in the influential article What is Web 2.0 by Tim O’Reilly, so here I will limit myself to what directly concerns tags, or rather hashtags: folksonomy and what O’Reilly called “harnessing collective intelligence”.

Although it may seem unimaginable from today’s perspective, the platform of the Web was not always so interactive, responsive and open to users. This is understandable, though, for the creativity of website makers has always gone hand in hand with what technology allowed. For example, a web designer could not think, when designing a website, of the site having plenty of photographic material until 1993, when the Mosaic browser came onto the market; it was the first to allow graphic elements (e.g. photographs, graphic elements of the user interface) to be displayed not in a new window but directly inside the content of the web page. Mosaic can thus be considered the web browser that for the very first time allowed the Web to approach graphic design, which at the same time created a demand for web designers and information architects, that is, professionals who could shape these new and amazing graphic possibilities so that a website could give visitors information and experiences quickly and appropriately.

With the arrival of client-side scripting languages such as Javascript, of browsers, of the use of more suitable programming languages and databases, and above all of support for these features from the makers of web browsers, web projects emerged that allowed not only richer interactions but also the involvement of their users in creating web content. Whereas websites used to be static and served as a slightly more expensive substitute for printed company brochures (the theorist and designer Rachel Hinman actually called such websites brochureware), now websites became a dynamic space where the user was not just a visitor but also a co-author. This social revolution of the Web platform also meant that the Web finally began to make use of its greatest strengths: hypertext and hyperlinks. In his article O’Reilly notes that:

“Hyperlinking is the foundation of the web. As users add new content, and new sites, it is bound in to the [existing] structure of the web by other users discovering the content and linking to it. Much as synapses form in the brain, with associations becoming stronger through repetition or intensity, the web of connections grows organically as an output of the collective activity of all web users” (O’Reilly, 2005)

Among the prototypical representatives of the social revolution of the Web platform, O’Reilly counts, besides Wikipedia and Ebay, the once popular bookmarking project del.icio.us and Flickr, still active and now part of Yahoo. Both projects popularised the concept of “tagging”, that is, they allowed users to label the photographs or bookmarks they added with short words, the choice of which was entirely in the hands of the users, their creativity, education or knowledge. This user-defined classification of items earned the name folksonomy, which arose by combining the words folk (popular, of the people) and taxonomy.

Besides the fact that the social features of web projects and tags outsource the generation of content and its categorisation to users, thereby, among other things, democratising the Web as such, they also offer an extraordinary insight into the thinking of individual users, since when defining tags a user may well be influenced by current conventions and the popularity of certain tags, but otherwise, as a rule, on today’s popular web projects that allow (hash)tagging nothing in theory prevents him from tagging as he pleases. In practice, however, the user does choose how to tag his photos or messages so as to make the best use of the whole essence of tagging: a brief comment on what the item is about, what it depicts. Certain tags can become so popular that they create a kind of coherent story; this can often be seen on today’s social networks such as Twitter, which dynamically builds a list of the currently most used tags. Their popularity often corresponds to some significant socio-cultural event, such as the announcement of the Oscar winners, a political scandal, a war or a natural disaster. Such a use of (hash)tags, however, already describes the present situation: Twitter did not offer (hash)tags right after its launch, and when it introduced them it assumed they would be used somewhat differently.

Twitter and the hashtag

When Twitter was launched in 2006, it offered “no technical or social mechanism for replying to another user, grouping tweets together, or indicating that a tweet was part of a wider topic.” (Highfield, 2015) It was only in 2007 that Chris Messina, at the time a Google employee, proposed a coherent idea of how Twitter could use the “#” sign to group Twitter messages. Messina drew inspiration from the IRC (Internet Relay Chat) protocol and chat, where the hash symbol is used to mark channels and topics, that is, for the same function Messina wanted to introduce on Twitter. One of the reasons the hashtag (hash + tag) caught on quickly was that it required few changes to Twitter’s existing infrastructure; moreover, the hashtag did not require users to have any technical knowledge of coding or searching. (Messina, 2007)

The introduction of the hashtag thus originally served to group Twitter messages by the same topic. (Scott, 2015) Some academics studying social networks, however, noticed that as early as between 2009 and 2010 this purist use of hashtags quickly “went rogue”. Others claim that hashtags stopped serving to organise content and turned into a “linguistic tumour”. (Vosper, 2016)

In other words, the hashtag changed from a utilitarian tool into a tool with which users began to express their emotions, without considering whether such a personal and emotional hashtag would contribute to the categorisation of the Twitter post, or to its better visibility and findability.

What an analysis of hashtags on the social network Twitter must, in my opinion, inherently assume is the fact that whether a hashtag stands for subjective emotional relations or, on the contrary, for an objective and deliberate categorisation of the post for the sake of easier searching, the chosen hashtag is necessarily related both to the post proper and to the other tags within the same post.

If that is so, I do not consider hashtags that stand for emotions a problem. On the contrary, an analysis of hashtags can uncover not only related topics (e.g. when topics are indicated by hashtags used within one post) but also what personal attitude and emotions users hold towards a given topic (e.g. when a user in one post uses one of the hashtags to stand for the main topic and one or more hashtags to express a personal, emotional attitude to that topic).

If I return to the original topic of this text — the analysis of the US presidential election — I plan to use the hypotheses mentioned above: if a Twitter post contains several hashtags, these hashtags will represent, among other things, the relations between the main topics of the particular post proper and the relation between the main topics and subjective, emotional reactions. Given a sufficient number of posts, the analysis should uncover the structure of a network of relations of the type topic–topic and topic–emotional reaction.

Method

For my analysis I chose four presidential candidates, two running for the Democratic Party and two for the Republicans. They are Hillary Clinton, Bernie Sanders, Donald Trump and Ted Cruz. For each candidate I chose suitable hashtags that are significantly associated with the candidates’ presidential campaigns. Such an analysis is not exhaustive, since there are several popular hashtags for any one candidate, but for a pilot study I consider the chosen hashtags sufficient.

I obtained the data for the chosen hashtags using the web application http://socioviz.net/, for which a free user account is available. It is limited by how many results the application returns to the user for one query. At present, users with a limited free account are allowed to get up to 100 Twitter posts per query from a very limited time interval of one day. This annoying limit can be partly circumvented if the user downloads the data separately for individual days. Unfortunately, even then the limited time interval cannot be changed, so the Twitter posts for a given day will come from a one-minute interval between 18:58 and 18:59. For my analysis I chose data from 1 April 2016 to 6 April 2016 inclusive. Altogether, I thus worked with 600 Twitter posts for each hashtag.

Chosen hashtags

  • Hillary Clinton: #iamwithher
  • Bernie Sanders: #feelthebern
  • Donald Trump: #donaldtrump
  • Ted Cruz: #tedcruz

Gephi settings

  • Layout: ForceAtlas 2
    • Dissuade Hubs
    • LinLog mode
    • Prevent Overlap
    • Edge Weight Influence 1.0
    • Scaling 2.0
    • Gravity 0.2
  • Degree Range filter: 5 (12 for #tedcruz)
  • Modularity Class

Gephi settings for the “tag cloud”

  • Layout: Fruchterman Reingold
  • Degree Range filter: 12
  • Show edges: No

Results

For each candidate I exported two types of graphs: one graph in the ForceAtlas 2 layout and the other in Fruchterman Reingold, which seems clearer if we can display the hashtags as a “tag cloud”, without an emphasised distance from the central node. For the “tag cloud” view I switched off the edges and, in addition, raised the lower limit of the filter for the total degree of a node. Let me recall that in a directed graph, where edges have a direction, which is my case, the total degree k of node i is calculated as the sum of incoming and outgoing edges, expressed mathematically as:

A14maT-Bc65dZyqW28yvhuy_Ovzk8UDcfdqZk_IpgBD-1iSzfFUYifoPK3zmQVt-LdeXa8JRwqOTiQ5t6BnJV4ItdYjkFvgPZvL3yjIY58pEjpFE3Gi04CL7ij-TT7ktQEDKEi_d

Finally, for each candidate I give a table of the most used hashtags that users used together with the primary hashtag.

Hillary Clinton

#iamwithher

#iamwithher hashtag, forceatlas2, degree filter 5 (click for full resolution 8192 x 4096)

#iamwithher_tagcloud

#iamwithher hashtag, tag cloud, degree filter 12 (click for full resolution 8192 x 4096)

[embeddoc url=“http://www.jakubferenc.cz/wordpress/wp-content/uploads/2016/04/iamwithher-Nodes-1.xlsx“ download=“all“ viewer=“microsoft“]

#iamwithher hashtag, tag cloud, degree filter 12, table of hashtags co-occurring with the primary one, sorted by weight

Bernie Sanders

#feelthebern

#feelthebern hashtag, forceatlas2, degree filter 5 (click for full resolution 8192 x 4096)

#feelthebern_tagcloud

#feelthebern hashtag, tag cloud, degree filter 12 (click for full resolution 8192 x 4096)

[embeddoc url=“http://www.jakubferenc.cz/wordpress/wp-content/uploads/2016/04/feelthebern-Nodes.xlsx“ download=“all“ viewer=“microsoft“]

#feelthebern hashtag, tag cloud, degree filter 12, table of hashtags co-occurring with the primary one, sorted by weight

Donald Trump

#donaldtrump

#donaldtrump hashtag, forceatlas2, degree filter 5 (click for full resolution 8192 x 4096)

#donaldtrump_tagcloud

#donaldtrump hashtag, tag cloud, degree filter 12 (click for full resolution 8192 x 4096)

[embeddoc url=“http://www.jakubferenc.cz/wordpress/wp-content/uploads/2016/04/donaldtrump-Nodes.xlsx“ download=“all“ viewer=“microsoft“]

#donaldtrump hashtag, tag cloud, degree filter 12, table of hashtags co-occurring with the primary one, sorted by weight

Ted Cruz

#tedcruz

#tedcruz hashtag, forceatlas2, degree filter 5 (click for full resolution 8192 x 4096)

#tedcruz_tagcloud

#tedcruz hashtag, tag cloud, degree filter 12 (click for full resolution 8192 x 4096)

[embeddoc url=“http://www.jakubferenc.cz/wordpress/wp-content/uploads/2016/04/tedcruz-Nodes.xlsx“ download=“all“ viewer=“microsoft“]

#tedcruz hashtag, tag cloud, degree filter 12, table of hashtags co-occurring with the primary one, sorted by weight

Commentary on the data

coming soon

References

HIGHFIELD, Tim and Tama LEAVER. 2015. A methodology for mapping Instagram hashtags. First Monday. 20(1), -. DOI: 10.5210/fm.v20i1.5563. ISSN 13960466. Also available at: http://journals.uic.edu/ojs/index.php/fm/article/view/5563

MESSINA, Chris. 2007. Groups for Twitter: or a proposal for Twitter tag channels [online]. [cited 2016-04-07]. Available at: http://factoryjoe.com/2007/08/25/groups-for-twitter-or-a-proposal-for-twitter-tag-channels/

O’REILLY, Tim. 2005. What Is Web 2.0: Design Patterns and Business Models for the Next Generation of Software.O’Reilly.com [online]. [cited 2016-04-07]. Available at: http://www.oreilly.com/pub/a/web2/archive/what-is-web-20.html

SCOTT, Kate. 2015. The pragmatics of hashtags: Inference and conversational style on Twitter. Journal of Pragmatics. 81, 8-20. DOI: 10.1016/j.pragma.2015.03.015. ISSN 03782166. Also available at: http://linkinghub.elsevier.com/retrieve/pii/S037821661500096X

VOSPER, Yuwa. 2016. Hashtags: Not Just Used in Social Media [online]. Louisiana State University Baton Rouge, LA [cited 2016-04-07]. Available at: https://www.academia.edu/23491466/Hashtags_Not_Just_Used_in_Social_Media