I was ready in bed. She lay obediently on the mattress next to mine, listening for whether she might hear me say her name. That day I had needed many things, but she had always rushed to help me with an answer or willingly met my various needs. My next desire, however, was to decide whether she would remain my intimate companion, or whether I would start looking for another one from the rival businesses of Google, Microsoft or Amazon. Her name was Siri, she is the digital assistant in Apple’s smart devices, and I had just, in all seriousness, asked her to bring me a chilled bottle of beer from the fridge.

Although the reader may read the opening paragraph as an excerpt from a poor third-rate novel, or as a picture of an unambiguous abuse of power between a man and a (digital) woman, in reality it is I, the user, who was abused. Without my consent, the technology of the digital assistant got into my head, changed something in it, and the result of this process was that I wanted to utter, for the first time in my life, a sentence that I would never before have said with the same seriousness and the same trust in the expected result as now. I wanted my phone, or rather its software part, the digital assistant Siri, to bring me something to drink from the fridge. How did I get there?

It all began when, after many years, an Apple smartphone once again became part of my life. Long ago I was one of the first people around me to get an iPhone 3G, thanks to a good discount from T-Mobile, which was offering the revolutionary Apple phone at an attractive price with a monthly plan. Since then I have gone through several phones and operating systems. For the last few years, however, I had owned a phone with Microsoft’s operating system, hoping that this technology giant would pull itself together and start releasing interesting apps on its platform. Over time I gave up this hope and made do with the expectation that Microsoft would at least keep the system viable. But in the end I did not live to see even that, because in July 2017 Microsoft officially ended support for its mobile operating system Windows Phone[1]. I took to the new Apple phone quickly; unlike the Nokia with Windows Phone, where with growing intensity almost nothing worked, on the iPhone with its finely tuned iOS 12 everything works so far. And one of the new attractions that I could not resist trying was the digital assistant Siri, whom I had heard about only second-hand and whom I now had the chance to interact with first-hand.

I started fairly conventionally. I asked Siri to play my favourite songs from a playlist, to set an alarm for the next morning, to activate a daily reminder to check all my emails at eight in the morning, etc. The interaction was not without problems, but Siri always tried to accommodate me as best she could. If she was not sure about the content of my request, she politely asked me to repeat once more what exactly I wanted from her. Because the new iPhone represented a real technological leap for me, my enthusiasm for discovering new functions and ways of making routine tasks easier did not wane in the days that followed either. One evening I was getting ready for bed. I was thinking about which book by the historian Timothy Snyder I should get. I opened the Safari browser on the iPhone and started googling information about the individual books. I decided to read in bed until I fell asleep. I lay down in bed, and just then I realised how handy it would be to have something to drink beside me while reading. Because in recent days my mind had been trained to use the option of faithful Siri whenever I needed to automate the bigger or smaller trifles of life, this time too I did not hesitate. I formed a mental image of the fridge in the kitchen and, without hesitation, automatically uttered the unthinkable: “Hey Siri, bring me a beer from the fridge”.

It took me about a second to realise that Siri would not be able to bring me a ten-degree Krahulík from the Czech family brewery Zichovec. But in the very next second I began to be more and more fascinated by the moment that had just taken place. How is it possible that I asked my phone to do something for me that is objectively impossible, given that Siri, so far, has neither legs nor an online connection to my unsmart fridge?

The first explanation that suggests itself is that I was simply tired and was not paying much attention to what I was saying. But I do not consider that an explanation. It is on the same level of argumentative sophistication as if we said that the human race created technology only because it was tired of doing routine tasks. Or might such an explanation, unintentionally, hide some truth within it?

I thought about this situation in terms of having taught my brain laziness during the days spent delegating tasks to Siri. I cannot say how exactly our individual routine matters are organised and stored in the brain, but is it possible that by using Siri I began to redesign these existing neural networks? Surely it is not hard to imagine that until then I had neural networks in my brain that solved the problem of how to turn on music on my phone, or what it takes to go to the fridge for the coveted object. My brain was building a relationship between the desired goal of what I want to achieve (a chilled beer) and the psychophysical reaction of how I will achieve that goal, in such a way that my self would consider the goal accomplished. This way of learning – action, reaction and evaluation of the result, followed by an adjustment of the original state of my self – is well known in today’s world of machine learning as learning by feedback (reinforcement learning). At its core lies a simple idea: at the beginning, the (software) agent has a goal; it can achieve it by means of some activity; once the activity has been carried out, it is evaluated whether the activity actually led to the desired goal.

Figure 1 A model of reinforcement learning

This model is an abstraction of the complexity of how we actually solve problems. It simplifies all the contextual aspects of the time and place in which we find ourselves, including the sociocultural influences that serve as regulators of goals and of the activities that come into consideration in a given context. It is precisely how our problem-solving involves ad hoc improvisation that the anthropologist and HCI theorist Lucy Suchman (1987) writes about.

Without Siri, under normal circumstances I would, for example, set the alarm on my phone myself. Such an activity would consist of several successive steps leading to my opening the Clock app, finding the alarm settings in it, entering the desired time and finally confirming my choice. The mental map, or rather the mental model, of the activity leading to setting an alarm on a smartphone is uncomplicated. But that is more a historical accident than an absolute truth. If the same goal were to be accomplished by a person who has never seen a smartphone because they live in one of the last remaining indigenous communities on this planet, this goal, so simple for us, would suddenly be enormously complicated for that person, indeed even impossible. For an indigenous person has never had a place for such goals in their conceptualisation of the world, because nothing like it exists in their world. From this point of view, technology appears to me as a medium that extends the hypothetical mental space of the possible, of what we may or may not think about.

On the other hand, even for us this activity can appear more complicated if we find ourselves in a situation where the smartphone is not simply lying on the bedside table but has, for example, a dead battery or some other fault that makes it harder to use. In such a case there are several ways I can reach my goal. Do I put the phone on the charger and wait until it charges to some reasonable level, or, as soon as I plug it in, do I try to turn the phone on, since I need to enter the alarm setting quickly? Because the space of potential activities leading to the desired goal holds several of them, it seems inevitable to me that there exist in our mind conscious, but also unconscious, selectors, which choose which activities are most suitable in a given context for achieving the set goal.

How do selectors work? Let us imagine that each of our goals stands in a relationship with a mental mechanism that chooses for us a suitable activity ensuring the goal is reached. Let us call this mental mechanism precisely a selector. Of course, one goal can be reached by many different activities. That is why the selector must take care of evaluating how well, on the basis of past experience, a given activity has proven itself in fulfilling the goal. However the evaluation of success in fulfilling goals actually takes place (the brain’s system of rewards, the so-called “reward system”, will play a role in it), the selector assigns the activity either plus or minus points, which indicate how suitable the activity is for fulfilling the given goal. In Figure 2 we see such a depiction, based on the previous diagram

Figure 2 A modified model of reinforcement learning using a selector

More astute readers are surely asking whether it is not possible to use more than one selector. My answer is that it is possible, and that there may be several selectors which are themselves chosen and evaluated at the hierarchically higher level of meta-selectors. But who or what chooses among the meta-selectors? This hierarchy of selectors cannot be infinite. Which is a problem that artificial intelligence (AI) researchers ran into in the seventies and eighties when, within the older paradigm of symbolic artificial intelligence, they tried to solve the problem of evaluating the relevance of subsequent actions. Marvin Minsky introduced into AI research the theory of so-called frames[2], which solves the problem of relevance by having the artificial intelligence system evaluate subsequent states on the basis of data structures – frames – which, in a symbolic, language-like form, represent formalised stereotypical situations from real life and contain relevant factual data and actions, explicitly defined by programmers, that “make sense” in the context of the frame. So, for example, within a frame named “kitchen” it makes sense to cook, to wash the dishes, a carrot, a knife and baking cakes, whereas the action “go to the toilet” or the object “Boeing 747” make somewhat less sense.

This “symbolic” conception of artificial intelligence, however, turned out to be too simplistic. Although it could solve logical and mathematical tasks such as playing chess reasonably well, tasks such as moving about and perceiving objects in the real world remained too complicated. Moreover, the problem of relevance, which frame theory was supposed to solve, persisted. As the philosopher and phenomenologist Hubert Dreyfus pointed out, frame theory merely moves the problem of relevance one step higher. An artificial intelligence system could indeed, for example, use the frame “kitchen” to determine roughly what appears, before the mechanical eyes of the system, to be relevant for the next steps in solving a particular problem. But Dreyfus pointed out (Dreyfus, 1979; Dreyfus, 1992) that the system for selecting the appropriate frame again needs a superior mechanism to choose the frame; in other words, some meta-frame is needed.

To avoid an infinite regress, Dreyfus proposes a solution inspired by phenomenology: the frame problem, as Dreyfus called it, is a much deeper problem than it may seem at first glance. Since it is out of the question that we could cram into an artificial intelligence system, in advance, all the facts, actions and frames about the world declared in a formal language, Dreyfus argues that it is our lived experience, our being-in-the-world, that ultimately defines what has meaning and relevance for us in the world in a given scene and what does not. Here Dreyfus builds on Heidegger’s concept of Dasein by accepting the thesis that human being is historically and socioculturally bounded. It is culture, society and their development that shape the relevance of objects and of the world around us. If artificial intelligence is to behave like a human being (and to solve the frame problem), it needs to have the same, or at least similar, experiences of the world as we humans have, and the world itself must in many cases become its own “best possible representation” (Brooks, 1991): without having to encode everything around us into some formalised symbolic language, we use the world as our representation, and in many cases we do not notice all the parts of the world but only those that provide us with affordances. By the word affordance I mean here above all the perceivable functions and possibilities for action that the environment around us provides through its appearance, form, material, etc. Some affordances are almost universally understandable across cultures; for example, a door handle offers the affordance of opening the door. Other affordances of objects, or in other words their functions and meanings, are, however, to a great extent shaped by culture and by the context in which we currently find ourselves. My favourite example is the PET bottle, which obviously offers the function of a container for liquid and the meaning of wasted plastic, but which, in a suitable context, can at the same time also be a flowerpot in ecological gardening, with a completely opposite and positive meaning.

Dreyfus’s argument, strongly influenced by Heidegger’s Dasein, says that what prevents the infinite regress of frames and meta-frames are the sociocultural layers and the physical world itself, because they create a structural limit on what is personally relevant and meaningful for us in a given situation. This does not mean, however, that these influences act, to use the previous diagram, only at the level of activities or selectors. They certainly act at these levels, but socioculture and the environment also define our goals themselves. Of course, some of our most basic goals, such as surviving, eating and drinking, are formed by a biological layer not mentioned so far; other goals, such as “buy a wedding ring”, are more or less sociocultural, and whether we buy the ring in Prague or in Ostrava is where the “spatiotemporal” layer comes into play, the one that most directs our goals and suitable activities in the context of the “here and now”. The previous diagram, modified to depict all the layers mentioned, could look like Figure 3.

Figure 3 A modified model of reinforcement learning showing the layers and their mutual interaction and co-constitution

What does the current form of the diagram show us? Let me remind you that the content in the middle, from the goal through the selector and the activities to the result, is an abstraction of a mechanism that we find implemented in some form not only in reinforcement learning within the current paradigm of artificial intelligence research, but also in the biological substrate of our brains, which give rise to our mind and to our cognitive skills, problem-solving among them. Let us note that although the individual layers are separated in the diagram, in reality they exist in mutual interaction and co-constitution, which makes them far more intertwined layers of our reality.

But in which layer is technology to be found, and where do we find an explanation that would shed light on why I suddenly wanted Siri to do something that I immediately judged to be nonsensical?

At first glance, technology fits neither into the biological nor into the spatiotemporal layer. For this reason it makes sense to assign technology to the sociocultural layer. That seems a relatively correct place to look for technology. However, given the semantic baggage this concept carries, it will be better first to take a step aside and define technology as something closely tied to socioculture, which nevertheless brings something of its own and something new that cannot simply be integrated into the existing conception of the social and the cultural.

The social, as studied by sociology, is concerned above all with human actors and their mutual relationships, which may or may not give rise to larger wholes in the form of families, ethnic groups, nations, etc. By the cultural I understand above all the ideas and objects that fit into one of the many categories of artistic genres.

By technology, then, I understand human-made objects, practices and processes which, at the beginning of their life, come into being above all as tools for solving concrete problems. But because technology does not exist in a vacuum, but is rather an integral part of our social and cultural world, I fully agree that technology also carries with it another layer, one that is shaped by socioculture. But the influences of socioculture have limits. For example, we may wish to have faster internet on our smart mobile devices wherever we go, yet the physical laws governing the propagation of waves of the electromagnetic spectrum set fixed limits independent of the sociocultural context. In short, we may passionately wish that our internet connection were much faster, but in the end we will always run into the limits of the speed of light, which sets the upper bound on the speed at which information propagates in our universe.

Technology is thus not an amorphous mass that we can knead into the desired form at will; rather, it has the power to resist us thanks to its objective materiality, which does not disappear even when we change the language of description in which we speak about the given material object. On the other hand, the functions and meanings of technology are not purely part of technology itself, but are created as the result of human interaction with technology in particular spatiotemporal contexts. Because technology influences socioculture just as socioculture influences technology, it seems to make no sense to depict their mutual relationship as hierarchical, with one side having more power than the other. That is also why the deterministic views of technology and socioculture still popular today – among them technological determinism and radical social constructivism – appear to me as views that always see only one side of the coin and refuse to look at the other. More appealing, and perhaps also closer to the truth, is an approach to this relationship that resembles the relationship in which the other layers stand: a closely interconnected and interactive one. On top of all this, we must take into account Heidegger’s contribution to thinking about technology, for in his well-known essay The Question Concerning Technology he points out that technology is not merely the result of applying scientific knowledge; rather, technological artefacts have always been an integral part of scientific progress: astronomy could not have done without the telescope that Galileo Galilei improved, just as computing simulations of fluid mechanics could not do without the brute force of today’s most powerful supercomputers. If we consider science a part of the culture of society, then the diagram of the relationship to technology would look like Figure 4.

Figure 4 A close-up of the relationship between socioculture and technology

In my diagram, technology becomes an integral part of this world and, like socioculture, shapes our thinking. At first glance this may seem a strange claim, but only until we abandon the long-held view of technology as passive pieces of matter that do nothing by themselves and wait until some human being (a subject) starts using them and turns them into a thing for use (an object). Instead, we should understand every introduction of a new technology as shaking up the technological background against which we act in the world, because it brings with it new technologically mediated affordances that change our psychological as well as cultural possibilities.

The view of technologies as a medium is, after all, nothing new. One of the best-known media theorists, the Canadian Marshall McLuhan (McLuhan, 1991), the French sociologist Bruno Latour (Latour, 1994), and today the philosophers of technology Don Ihde (Ihde, 1990) and Peter-Paul Verbeek (Verbeek, 2016) or the media theorist Mark Deuze (Deuze, 2015) all consider information and communication technologies to be synonymous with media.

This step is not purely rhetorical but has far-reaching consequences, for it essentially defines a new metaphysics of the relationship between human beings and technology. After McLuhan, Latour, Ihde and Verbeek, we know that technology as a medium is not neutral and influences both our actions and our relationship to the world. This means that in our diagram we must look at goals and activities not as given in advance, with technology passively carrying out our will, but rather as an imaginary space of potentialities into which technology, through its influence, inserts new potential goals and new activities for achieving these goals.

Latour’s favourite example of technological mediation is the firearm. He argues that the gun by itself does not kill, just as it is not the human alone who kills. He refers to the well-known slogans used in the USA in the long-running, passionate debate over gun control. Liberals like to use the slogan “guns kill”, whereas the right-wing, Republican-leaning National Rifle Association (NRA), somewhat paradoxically, resorts to sociological analysis when it claims that it is “people who kill”, not guns. Latour argues that neither view is correct, for both refuse to admit that the other side’s argument might contain at least part of the truth. Latour argues that we should analyse a human being with a gun as a new entity, the human-gun, which has new attributes and the potential for new activities that we will not find if we analyse the human or the gun separately. The power to kill and the responsibility for killing are thus distributed between the human and the gun.

My favourite example is the sketchbook and pencil that an artistically gifted person works with. When a person sketches, the resulting drawing is rarely similar to what they imagined at the start. The drawing is thus not stored complete in the artist’s head, with the artist then spending minutes or hours transferring the already finished drawing onto paper point by point. Instead, as Andy Clark notes (Clark, 1997), sketching can be understood as a continuous interaction of the person, the pencil and the sketchbook, in which the resulting drawing emerges because the person has constant feedback on the state of the drawing, provided by the sketchbook. Without this feedback, the artist would not complete the drawing. We can understand this as the sketchbook being an extended and externalised working memory from which the artist draws information for the next steps. Because the sketchbook plays such a crucial role in completing the drawing, we should speak of the authorship of the resulting drawing as being distributed between the artist and his technology (the sketchbook as well as the pencil). In the vocabulary of cybernetics, the resulting drawing is an emergent output of the interaction of the components of the newly formed human-sketchbook-pencil system.

So what about Siri?

When I started using an iPhone and a digital assistant in my life, my life and my skills for solving everyday problems were enriched by new possible goals and activities that I can pursue. For some goals Siri acts as a medium, as in the example of turning up the music, which I could also do with the mechanical buttons on the side of the phone. Some goals are new thanks to Siri and the iPhone in general, for example replying to messages and comments in various social apps. The more I used Siri as a medium between my existing goal and the result, the more Siri became the natural choice for reaching my goals. In the diagram below, see Figure 5, I have added Siri as the dominant activity that mediates more and more of my goals. Unknowingly – and here lies the abuse from the first paragraph of this text – I was thus training my brain to gradually learn to automate Siri as a natural part of my cognitive resources for solving problems in the world.

Figure 5 The diagram modified by the previous changes, showing Siri as the dominant activity that mediates my goals

So far, all of this should make sense in the context of this text. My brain adapts to a new environment in which Siri exists as an activity-medium that solves my problems, and their successful solution stimulates the reward system in my brain, which assigns plus points to the activity “Siri”, thereby saturating the weighted value (in figure 3.7) to such an extent that the given selector never moves on to any activity other than Siri. Because its weighted value is saturated, the selector then de facto becomes the activity named Siri. My brain thus internalises Siri to such an extent that it no longer regards it as something external to my mind or cognition, and Siri becomes a natural part of it.

My speculative view is that an external artefact internalised in this way, one that has turned from a potential activity into a more or less automated selector in my mind, may cause my brain to start using this selector-activity named “Siri” even for goals that were previously not connected with the given selector and activities. What happens is something we might call, in psychology, a “transfer” of skills between domains.

The classic transfer of skills and knowledge in developmental and educational psychology I would call a “positive” transfer, for we can indeed trace the same skill in another domain. Yet what kind of skill is involved, if any at all, in my case, when I suddenly wanted Siri to bring me a bottle of beer from the fridge?

I would call it a “negative” transfer, because the transfer of one internalised skill (solving problems with Siri) does not work at all in another domain. But just as in formal logic we can prove a solution directly as well as by contradiction, I will borrow the metaphor of proof by contradiction here to point out that even though I very quickly realised that Siri cannot bring me a beer from the fridge, my mind, extended by new possibilities, goals and activities, immediately allowed me to imagine what would be needed for Siri to be able to bring me a beer from the fridge. My idea was to get a smart fridge and a smart robot that can move around and has a digital assistant built into it. My concrete solution is, in the end, of no relevance. Far more interesting is the final result of my autoethnographic analysis of building a relationship with a new phone and the digital assistant Siri: thanks to the negative transfer of skills and knowledge, I was able to take an existing technology and, with a little imagination, picture a completely new way of solving my oh-so-pressing beer problem. Negative transfer opened up new worlds to me that I had not known until then I needed, or even that they existed.

 

Bibliography

BROOKS, Rodney A., 1991. Intelligence without representation. Artificial Intelligence [online]. 47(1-3), 139-159 [cited 2018-05-22]. DOI: 10.1016/0004-3702(91)90053-M. ISSN 00043702. Available at: http://linkinghub.elsevier.com/retrieve/pii/000437029190053M

CLARK, Andy, 1997. Being there: putting brain, body, and world together again. Cambridge, Mass.: MIT Press. ISBN 9780262032407.

DEUZE, Mark, 2015. Media life: Život v médiích. First Czech edition. Translated by Petra IZDNÁ. Praha: Univerzita Karlova v Praze, nakladatelství Karolinum. Studia nových médií. ISBN 978-80-246-2815-8.

DREYFUS, Hubert, 1992. What computers still can’t do: a critique of artificial reason. Cambridge, Mass: MIT Press. ISBN 9780262540674.

DREYFUS, Hubert L., 1979. What computers can’t do: the limits of artificial intelligence. Rev. ed. New York: Harper Colophon Books. ISBN 9780060906139.

IHDE, Don, 1990. Technology and the lifeworld: from garden to earth. Bloomington: Indiana University Press. ISBN 0253205603.

LATOUR, Bruno, 1994. On Technical Mediation. In: Common Knowledge. Vol.3 n°2. p. 29-64.

MCLUHAN, Marshall, 1991. Jak rozumět médiím: extenze člověka. 1st ed. Praha: Odeon. Eseje (Odeon). ISBN 80-207-0296-2.

SUCHMAN, Lucille Alice., 1987. Plans and situated actions: the problem of human-machine communication. New York: Cambridge University Press. ISBN 0521337399.

VERBEEK, P.P., 2016. Toward a Theory of Technological Mediation: A Program for Postphenomenological Research. In: FRIIS, Berg O. and Robert C. CREASE. Technoscience and Postphenomenology: The Manhattan Papers. London: Lexington Books, p. 189-204. ISBN 978-0-7391-8961-0.

Notes

[1] https://www.theverge.com/2017/7/11/15952654/microsoft-windows-phone-end-of-support

[2] https://en.wikipedia.org/wiki/Frame_(artificial_intelligence)