{"id":2523,"date":"2025-11-03T08:21:50","date_gmt":"2025-11-03T08:21:50","guid":{"rendered":"https:\/\/blogs.shu.ac.uk\/sioe\/?p=2523"},"modified":"2025-11-03T08:21:50","modified_gmt":"2025-11-03T08:21:50","slug":"think-before-you-chat-gen-ai-and-applied-linguistics","status":"publish","type":"post","link":"https:\/\/blogs.shu.ac.uk\/sioe\/2025\/11\/03\/think-before-you-chat-gen-ai-and-applied-linguistics\/","title":{"rendered":"Think before you chat: Gen AI and applied linguistics"},"content":{"rendered":"<p>\u201cYou shall know a word by the company it keeps\u201d (J.R. Firth, 1957)<\/p>\n<p>Here at SIoE, we ask big questions about how people learn, teach, and communicate, and now that includes how machines are learning to \u2018chat\u2019 like us. Language is the beating heart of education, so it is not surprising that ChatBots, powered by generative AI, are causing palpitations. With generative AI shaking up classrooms, assessments, and everyday communication, a look at the linguistic roots of large language models can tell us more about the possibilities and the limitations of this technology.<\/p>\n<p>Have you ever asked ChatGPT, Gemini or Jen.AI a question and thought, <em>\u201cWait\u2026 did a bot really write that?\u201d<\/em> You&#8217;re not alone. These days, bots don\u2019t just chat, they charm, explain and even argue like the best of us. When you read an AI-generated answer, do you believe that a human with access to the same information may have written those lines? If so, it means that Generative AI has passed the \u201cTuring Test\u201d. This long-standing benchmark for machine intelligence \u2013 that a machine can trick a human into thinking they are interacting with another human \u2013 has been swept away thanks to large language models. While the furore in the media shows this is a surprise to many, applied linguists have been waiting for this day since the 1950s when Firth, the UK\u2019s first professor of linguistics, announced that \u201cYou shall know a word by the company it keeps.\u201d<\/p>\n<p>Linguists have long known that there is something mathematical (and magical) about language. Since 1932, \u201cZipf\u2019s law\u201d has told us that a word\u2019s frequency in text is inversely proportional to its rank, so that around 1,500 of the most common words (starting with <em>the<\/em>, <em>of<\/em>, <em>and<\/em>, <em>to<\/em>) account for about 80% of all English text. We know that, overall, any body of texts shows the same pattern: a small number of highly frequent words (or lexical items) do most of the heavy grammatical lifting in language, while a very large number of low frequency lexical items specify topic areas. We can evaluate evidence for Zipf\u2019s law because, since the technology became available in the 1980s, corpus linguistics has analysed increasingly large bodies of machine-readable language, expanded exponentially in the last three decades by the proliferation of texts on the internet. Corpus linguists use the search and pattern-matching abilities of computers to find grammatical and lexical (or lexico-grammatical) patterns. By examining language empirically, it is corpus linguistics that laid the theoretical and practical foundations for the large language models that generative AI depends on.<\/p>\n<p>Corpus studies have provided the data for words that, as Frith predicted in the 1950s, can be known through knowing the accompanying words that regularly share the same locations \u2013 their collocations. Words that appear to be very similar in meaning (like <em>start<\/em> or <em>begin<\/em>) will invariably collocate with different sets of words (Sinclair, 1991). Words that have multiple senses (like <em>post<\/em> and <em>wear<\/em>) will collocate with different words in their different meanings. However, the location of a word may be even more consequential than previously expected \u2013 it seems that words are also \u2018primed\u2019 to appear at the start, middle or end of a sentence, a paragraph or a text (Hoey, 2004). In fact, it is possible to predict, with a large enough dataset of texts (known as a corpus), the probability of any two words appearing together. This is the principle that drives generative AI \u201cspeech\u201d.<\/p>\n<p>While this may sound, computationally, very straightforward, machines struggle with their predictions because it is the social context of the text that decides, when collocations are equally likely, which collocate should be chosen. Words keep the company of other words, but Firth also discussed how they keep the company of a social context. Generative AI does not <em>know<\/em> the difference between arguing with a taxi driver and consoling a grandparent; they do not understand Register. Register is also magically mathematical \u2013 probabilities in lexico-grammar vary according to social relations, the topic and the role of the text \u2013 but this is exactly the point where Generative AI fails. It is only with human training that the AI bots can provide a response that a human will consider meaningful based on their contextualising input, and it is here that the tech companies spent big on humans to train their AI bots to understand what is contextually appropriate. Without this very human input, generative AI remains totally ignorant of the human experience.<\/p>\n<p>By recombining our words, generative AI may dazzle us with its fluency, but it still lacks something fundamental: lived human context. It didn\u2019t <em>learn<\/em> language through experience and education. It can\u2019t <em>know<\/em> language because it doesn\u2019t know context; it can only calculate the most probable next word. While that may be enough to pass the Turing Test, it\u2019s not enough to understand meaning or intention, or to care, to love or inspire. As applied linguistics has long shown us, language isn\u2019t just about words, but about relationships: relationships between words, between people, and within society. Generative AI may have learned how to mimic those patterns, but it hasn\u2019t learned how to <em>live<\/em> them. This is exactly where our curiosity lies, not just in how language works, but in how it lives <em>in us<\/em> as humans, learners, and communicators. As Gen AI increases its hold on our languaging lives, it\u2019s our human relationships that will always remain beyond its grasp and will continue to separate human and bot chat.<\/p>\n<p>Firth, J.R. (1957) A synopsis of linguistic theory. In F.R. Palmer (Ed.) (1968) <em>Selected Papers of J.R. Firth 1952-59<\/em>. London: Longmans<\/p>\n<p>Hoey, M. (2004) <em>Lexical Priming: A New Theory of Words and Language<\/em>. Routledge.<\/p>\n<p>Sinclair, J.M. (1991) <em>Corpus Concordance Collocation<\/em>. Oxford University Press<\/p>\n<p>Dr Nick Moore is a senior lecturer in TESOL in the SIoE.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>\u201cYou shall know a word by the company it keeps\u201d (J.R. Firth, 1957) Here at SIoE, we ask big questions about how people learn, teach, and communicate, and now that includes how machines are learning to \u2018chat\u2019 like us. Language is the beating heart of education, so it is not surprising that ChatBots, powered by [&hellip;]<\/p>\n","protected":false},"author":10265,"featured_media":2524,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"wds_primary_category":0,"footnotes":"","jetpack_post_was_ever_published":false,"_links_to":"","_links_to_target":""},"categories":[4],"tags":[],"class_list":["post-2523","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-sioe"],"acf":[],"jetpack_featured_media_url":"https:\/\/i0.wp.com\/blogs.shu.ac.uk\/sioe\/files\/2025\/10\/typewriter-1248088_1280.jpg?fit=1280%2C853&ssl=1","jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p7axDP-EH","_links":{"self":[{"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/posts\/2523","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/comments?post=2523"}],"version-history":[{"count":4,"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/posts\/2523\/revisions"}],"predecessor-version":[{"id":2531,"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/posts\/2523\/revisions\/2531"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/media\/2524"}],"wp:attachment":[{"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/media?parent=2523"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/categories?post=2523"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.shu.ac.uk\/sioe\/wp-json\/wp\/v2\/tags?post=2523"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}