Humans in the lead: Beyond AI detection in education video
Hello world and welcome to the podcast for educators passionate about computing and AI.
I'm Rehana Al-Soltane, a learning manager here at the Raspberry Pi Foundation, and this is the third episode to accompany the latest issue of the Hello World magazine, which explores critical thinking in the age of AI.
In today's episode, I'm talking to Sam Illingworth, professor of critical AI literacy at Edinburgh Napier University and the founder of Slow AI, a newsletter and curriculum for critical AI literacy with over 19,000 subscribers.
Welcome to the podcast Sam!
Hi, Rehana. I'm absolutely delighted to be here.
Well, thank you so much for coming. I am one of your 19,000 subscribers on the slow AI newsletter, and whenever my inbox pings me that you have released a new blog and post, I get quite happy.
And that's mostly because I think critical perspectives on AI are quite hard to come by. And I think you have a very, very important critical perspective on AI technologies around us.
And I think that we live in a world. We exist in a world where we are encouraged to use AI technologies left and right, almost blindly, without truly understanding what impact it's going to have on our skills or who it was designed for, or understanding who it was not designed for.
And that's why I'm so excited to talk to you today, and to talk about critical AI literacy, and unpack that a little bit, and also talk about the article that you wrote for the Hello World magazine.
But before I dive into all these topics and all these questions that I have, I would love to hear more about you and what you do as slow AI and at the University. Great.
So I'm a full professor of critical AI Literacy at Edinburgh Napier University, and a lot of my work there involves looking at how students and staff are actually using AI tools and in particular, generative AI tools like ChatGPT, Claude, Gemini, etc.
and I'm the PI on the UK's largest qualitative research study, which was funded by the Leverhulme Trust and which we're just wrapping up at the moment.
And what that did was it surveyed over 7000 students in the UK to ask them how they're actually using these tools, rather than presupposing how they might be doing so?
So that's really a lot of what my work is around my, you know, my own background is I've PhD in atmospheric physics, but I'm also a poet. It's obviously a very natural combination.
And I think my interest in AI really came about with the release of ChatGPT in 2022. And when it came out, you know, I was like, oh, cool. This is an interesting tool to play with. I can use it to like make some haiku and stuff like that.
But what really struck me almost straight away was this is going to lead to the potential for an exasperation or a widening inning of the digital divide.
So, you know, the digital divide that we saw during Covid in schools with those students who had access to a laptop, to high speed internet, to software versus those that didn't.
So I was playing with these tools that this is really, really interesting. And a lot of my work after being a physicist was actually grounded in working with marginalised communities and listening to their voices and making sure that their voices were heard within the democratisation of science.
So for me, it was like, oh, it's really important to think about how this new technology is going to be adapted and adopted.
And the key thing is, right, these tools are mainly written by people who look like me white, heterosexual, cisgendered, Western men, and they're trained on the internet, which is dominated by white, heterosexual, cisgendered men.
So they're incredibly bias. They don't speak for all people.
My broad position in is there can be very useful, but critical AI literacy is really all about knowing when to use AI and knowing when to leave it alone.
And that's really in a nutshell, what my work is all about. It's about working with different communities, be it staff, be it student, be it other groups, and really thinking about how they might use these AI tools, and also, sadly, how sometimes those AI tools might use them as well.
Yeah, exactly. I find it so interesting that your background is so multidisciplinary. I think when you have a multidisciplinary lens and you look at AI technologies around you, I think you can see those patterns, like you said, about some communities not having access to begin with, so that some of these communities don't have access to these technologies to start with.
And that's the that's the divide that we're really talking about. But also then taking that to then advise those communities and how to use them responsibly and maybe sometimes not to use them as well.
Recently authored an article for the Hello World magazine on why AI detection software tools in education are not really that good. They're not the best use case, and that we shouldn't really use them in the field of education.
I agree with you 100%, but I would really like for you to explain why you think AI detection software tools don't really work as intended. In the world of education.
I try to be really central with my approach for most things related to AI and to genuinely take this central approach of there's no point banning AI, there's no point adopting AI piecemeal, but rather to have this central critical ground.
However, the one area that I don't hedge and that I feel very strongly about is AI detection.
So I'll explain why AI detection will never work.
Firstly, purely technical level, it will always bias certain people.
So how does AI work? Well, AI how does generative AI work? Generative AI works by basically predicting the most likely next token to appear.
So what that does is it's trained on the corpus or change on the large body of work, and it predicts the most likely sentence, the most likely word to come next.
How do AI detection tools work? Well, what they do is they look to see if the words and sentences that have been constructed have a strong predictability.
Now, the problem with that is that for anybody who's learned English as a second language or any other language, as a second language, as a non-native speaker, you'll know that you speak in a much more formal, much more structured way.
Similarly, for people that are neurodiverse, you communicate in a way that is more structured and that has more of a pattern to it than somebody who might identify as being neurotypical or a native speaker.
So what this means is that these tools, even if they work correctly, which they don't, often flag non-native English speakers and neurodiverse people the most.
There was a brilliant paper in 2023 from Stanford, which basically got a load of student essays written by humans and put them through AI detectors.
And for native English speakers, about 5% were falsely flagged as written by AI. For non-native English speakers, it was over 90%.
Now, this was three years ago, so the technology has definitely improved.
But it's fascinating because we're talking about Substack. Substack just literally this week signed an agreement with Pangram, which is another AI detection tool, software, so it can run your your newsletter through AI.
Now I use AI to help with my editing and my research and everything, and what was really interesting was when I, when I got my post and then put it through a humanised, which is another AI tool.
Pangram told me that it was 100% written by a human, whereas at that point it had to be written 0% by a human.
So we know that these tools don't work.
But also Rehana, even if they did work perfectly, even if they didn't prejudice people who have already been prejudiced, it creates an illusion of mistrust.
And in the classroom, what that tells me is if I'm using AI detection software, I'm instantly telling my students, I don't trust you.
And that's not why I got into education. I got into education for pedagogy, not policing.
So what I do is instead I say, look, we can use AI tools, but let's be honest about it. Tell me how you've used it in the assessment. Talk me through these steps.
Tell me, is my assessment really boring or really hard or really non relevant?
So you've just used an AI tool to generate an answer to a 5000 word essay that nobody's going to read anyway. That's fine. Let's have that dialogue.
What I don't want to do is marginalise and penalise people who are already being marginalised. Just because they happen to speak and write in a way that's different from a native English speaker.
It's a very long answer. Sorry, but it's a soapbox I'm quite passionate about.
Well, I think it's actually very important to explain the whole context around it.
So you mentioned that non-native English speakers are disproportionately more often going to be flagged by these AI detection software because they use simpler grammar structures or simpler vocabulary.
And all this is then associated with AI slop or AI generated content.
I'm someone who's a non-native English speaker. I learned English as a fourth language growing up, and there is one one phrase that I that I enjoy using so often because there was an embodied metaphor and it was easy for me to learn and then to use.
And that phrase was delve deeper, because to me it meant you put on a swimsuit, you dive below the water deep into the water, and you explore what's what's underneath.
And it's a it's a kind of like a, it's a metaphor that embodies action. Right.
And I think when you're learning languages, the more you can associate with that, the easier it is to learn these phrases.
But now more and more AI generated content includes phrases like these.
And now delve deeper is actually associated with AI slop.
And for context, a lot of the frontier AI models have been trained using data from countries like Nigeria, where phrases like delve deeper are much more common than they are here.
And so just because that has been implemented, that has been, that is represented heavily in the training data. That means the output then is also more likely to include these phrases.
But that also means that we are now associating such phrases like delve deeper, but also other ones like rich tapestry and so many more with AI generated slop.
And that is quite unfortunate because I think I've been robbed of several of my favourite phrases from my daily vocabulary in my daily language.
So my question to you, Sam, is what do you think this is doing to the richness and the vastness of our language, and whether now we are incentivising students to avoid using expressive language, and that maybe in the long term, that means that our language use will become more uniform, maybe a little bit more boring, maybe a little bit less rich.
It's a great question. First, I want to acknowledge English as a fourth language is insane. I have a believable impressed as somebody who tried to learn Japanese whilst living in Japan and finding it incredibly difficult.
Yeah, yeah, I'm bowled over by that.
It helps when you're very young and.
It's still very impressive. I think you're right.
And you know, it's really important to acknowledge where those phrases come from.
I think also, you know, slops existed before AI, and for me, AI slop is not about certain words or phrases or patterns. It's about a complete lack of human judgement.
For one thing that's really, I mean, for me, grounded is a word I've always used and I can't really use any more because that always gets flagged as being AI slop.
The other thing is, you know, these structures that are incredibly powerful, like the antithesis, right?
You know, it's not X, it's Y is an argument that's been used for thousands of years.
And it's incredibly powerful. And it's because that's powerful. That's why AI tools have picked it up as an effective arguing tool. Right.
So I think sometimes we go a little bit too far as you've situated.
And what's really interesting Rehana is that there's actually research out at the moment that's showing that humans are starting to modify their language use to better mirror AI, which itself has changed based on human languages.
So we're caught in this, in this loop, right?
The other thing I think that I always get in trouble for saying this, but I will say it and I think that English is a very constipated language.
Like we can't actually say what we it's why we're repressed as a nation. We can't really say what we want to say.
If you look at something like Japanese, which is this beautiful, vast or Arabic language, there's all these words and ideas to express.
And then English, it's like I literally cannot communicate what I've tried to communicate with these words.
But we know that English is the lingua franca of science and of AI as well.
So not only are we kind of restricted to this language, which is already restricted we're then, as you say, being forced to some extent to remove all of those rich metaphors and, you know, cultural changes around that language use.
It's something we all need to be aware of.
And I think the biggest advice I would give is just instead of calling people out for using certain words or phrases or or writing tics, instead look at where that judgement is.
You know, it's fine to use AI to help you edit, to help you draft, to help you think of ideas as well.
But you need to be the person that's responsible.
Heard a great phrase this week: "We often heard about humans in the loop. Actually, it should be humans in the lead".
So this idea that we can work with those AI tools to draft our work, to edit our work, but we're the ones as humans that need to have the final decision on whether what we've created is something that we're willing to stand behind.
So this element of judgement that you just that you just explained, you also have included it in your definition of critical AI literacy.
So if I can read out what your definition of critical AI literacy is, which is "the ability to evaluate AI outputs, recognise their limitations, understand the systems behind them, and make informed decisions about when to use them and when not to."
I think that last part is actually very, very important.
If you could describe what an ideal critically AI literate student would look like, how would they look like? What skills would they have then?
That's a great question. I think for me, what they would have is they would have the ability to critically think without AI, first of all, and this is I mean, I genuinely try to stay free of hyperbole when talking about AI, but the thing I do worry about is a generation of students, again, to be experiencing education through the lens of AI. And people such as myself you know, I'm in my 40s, had decades of learning how to critically think before AI came out.
So in many ways, it's very easy for me to apply those skills to AI.
But what about those students who've never learned to critically think for themselves, and who are working towards an output rather than a process?
And part of that is the assessment issue. Assessment for me shouldn't be about monitoring outputs. It should be about assessing change and monitoring change.
So what I think we should do here is we should try to avoid something called never skilling, which is where students never get to learn those tasks in the first instance.
And a critically AI literate student, an ideal one is someone who wants to question things. He wants to talk about things, who understands their limitations as well.
And what that looks like in practice is requires modelling that behaviour as well.
And also, you know, I need to think about my own positionality. As a white Western man who's a full professor, it's very easy for me to say, I don't know.
And I often don't know because I'm in a position of responsibility and authority that other people don't have, that other people might not feel comfortable to say they don't know in that position.
So that's why I think people such as myself need to use their privileges to create these opportunities for dialogue and for meaningful, you know, exploration of this space.
So if we could increase more opportunities for dialogue, I think assessment is one of those opportunities that we could revamp a little bit and rethink what assessment is and essentially make it an opportunity for dialogue, make it into an opportunity where students can showcase their their learning and kind of celebrate their learning.
So I wanted to focus on the element of assessment, because you also mentioned that our way of assessing students right now is just designed in a way that focuses on output more than the process that led, that led to that output.
So if you could if you could think about an English classroom, let's say, and you were to redesign what assessment looked like for a typical English student, what would you suggest we change in terms of how can we help students learn to think better?
How can we help them to structure an argument or make up a hypothesis? How would you reimagine assessment if you could?
Exactly. So I can give a specific example like that. I think that does work really well because I do it.
And this addresses two things. It teaches critical eye literacy, but it also addresses that something that we call the hidden curriculum.
So this idea that a lot of educators understand how marking rubrics work understands how assessment works, but we assume our students to also understand that even though they might not.
So what I do, and I invite anybody listening to this podcast to do as well, if you're an educator, is set a normal English assessment task.
So let's say write me a 1000 words evaluation or like book report on this, on this, on this book. Write me a 1000 word book report, but don't actually get them to do it.
Instead, show them your marking rubrics. Ask them to put the prompt for the assessment into their AI tool of choice, and then to use your marking rubrics to mark the work that was outputted by the AI tool.
And what that does is it teaches the students where the strengths and where the limitations of that AI tool is, but it also gets them to understand how curriculum work and how to improve your marking matrix as well.
And one of the things that always comes out of this is that, you know, we talked earlier about AI tools and generative AI in particular. Generative AI predicts the most likely next outcome.
So what that means is that all of the, most of the responses you get are quite vanilla.
You know, they'll look at all of, you know, let's say we've got a book like 100 Years of Solitude by Gabriel Garcia Marquez, it's one of my favourite books.
What it will do is it will look of all, all of the thousands of articles that's been written about 100 Years of Solitude, and it will write a response that doesn't really offer anything new.
Whereas if a student instead read it and was like, reading this book reminds me of when I was a child, and my uncle tied himself to a tree in the courtyard for three hours as a joke.
Like, that's a very deep cut for that particular book.
Yeah, that's that's a much more interesting piece of work that an AI tool would never be able to come up with.
So that kind of assessment, that kind of work shows the limitations of these AI tools, but also highlights the strengths of, you know, the human mind in terms of creativity and originality of thought as well.
So really, I think this is a very interesting idea to kind of reverse engineer what assessment is and have the, and to have the students to be in the lead, like you said earlier as well, I think another element of improving how we assess students is also helping them understand how they learn and what it is that makes us learn, what kind of strategies help us learn, and what the meaning of learning is in general. Right?
Because I think at the moment, I think the way that our assessments work, or maybe our schools are structured, is that they're preparing children to be collaborating citizens of the future.
Oh my goodness, yeah. We're just putting them on a, on a what do you call that? It's like.
On a conveyor belt.
On a conveyor belt. Yes. So that we put them on a conveyor belt and Mary predestined destination.
I'm laughing as well, because basically. So I did a master's in physics with space science and technology. Then I did a PhD in atmospheric physics.
And then when I got my first tenured position, you know, I had to do a postgraduate certificate in academic practice, like learning how to teach.
And I turned that into a master's in higher education when I was like 30.
And it wasn't until then that I'd actually learned how to learn like, so I'd gone through I'd gone through the entire process of like seven years in higher education without ever learning how to learn.
And like that. Master's in higher education was so foundational for me for how things work, and I could not agree more in terms of assessment.
You know, there's a really great higher education researcher, Professor Jan MacArthur, and she talks about authentic assessment and authentic assessment not being, and the purpose of the university, not being preparing our students for the workforce, but rather to prepare our students to challenge the workforce and to challenge those societal issues and marginalisation and discrepancies.
And that's what assessment should be doing. Exactly. Like you say, it shouldn't be a conveyor belt. It should be encouraging students to answer back, to critique, to like, really express themselves in the fullest possible way.
And if AI tools can help them do that, which I honestly think they can, then that's something that we should be bringing into the assessment process, rather than completely removing and burying our heads in the sand.
So in that case, we really need students to be critically literate with or without AI, and to really strengthen their core skills that they had pre-AI. You mentioned as well that both you and I have had decades worth of education, so we've had the time to develop all these skills, to practice our critical thinking, to make up all these arguments or to structure arguments and structure entire essays. Right.
And so to then have students not have that opportunity, I think we're robbing them of that critical understanding, the critical skills that they need to be citizens of the world who can challenge the workforce, like you said, and challenge the status quo.
So I think that's very empowering that you said that.
I want to talk a little bit about some games that you developed. So there are two in particular that you also reference in the in the article that you wrote for Hello World, one is called Bot or Not and the other is called flagged.
And I had so much fun playing both of them. It was more eye opening I think, and hit to my ego if I, if I can be very honest.
And for context, the bots are not. Game is where you have to decide whether a content has been written by a human or has been AI generated.
And in the flagged game you get a score from an AI detection tool, and then you have to use that score to decide if you pass a student's essay as acceptable or you flag it for further inspection.
I played both of these games, and what I was most surprised about was how fragile my assumptions were around what is AI generated and what is human written.
And I think this is a very important concept, which is we aren't that good of a judges. We aren't that good at judging whether something has been AI generated or has been written by a human.
And I think that these games really convey that element of fragility in our, in our judgements, but also the uncertainty in our judgements. The kind of clues that we rely on are not reliable at all. That kind of ambiguous. They're pretty vague.
Why do you think it's important that we convey this sense of uncertainty and fragility in our judgement when we think about AI tools and AI detection software, for example?
And why do you think that highlighting this fragility in our judgement is such an important step towards critical AI literacy?
I think that highlighting those judgements is essential because otherwise what happens is it turns into some kind of witch hunt where we're not even, we're not even reading or engaging with the content.
We're just trying to think, was this generated by AI or not? Which I think is really dangerous now, you know, that's not to say that we shouldn't get go down one end of the spectrum, which is where certain social media sites have a lot of AI generated content, and then the comments of AI generated content and they lack that judgement.
That's not what we want. But on the other hand, I don't think there's an issue with people, as I've said, using AI for certain processes. I'm upfront about it.
I use AI to help me research. I use AI to help me improve certain turns of phrases.
But I checked everything myself. I read everything myself, and it's all my original thoughts. I mean, I don't particularly see a problem with that. I think it's important that we're transparent about it as well.
And you know, you didn't do that badly Rehana, I think the average score on or not is like 40%. So it's, it's really hard.
And the whole purpose of that game was just to demonstrate that you can't you can't tell like either through our detection or just by eyeballing something, whether it's been written by AI or not.
I wrote that game. There's like 100 quotes that get pulled randomly. I still have yet to beat nine out of ten, like some people have got 100%, which is fine, but I, I don't get it all the time because it's hard.
And the other game you mentioned flagged, you know, you can't win at that game, like the whole point of that game, is to show people that behind every student's assessment is a person and a human.
And even if a student did use AI to plagiarise a piece of work. Why did they do that? What's the reason behind that?
And do you just flag them? Or do you go and have a conversation and a discussion with them, or set up a learning environment in which they feel comfortable about having that discussion with you in the first place?
So for me, a lot of the work that we do and that I do around AI is about how we can really drive more human to human connection.
And I am a, I'm a staunch optimist, and I genuinely think that AI is creating an opportunity for humans to reconnect, because I think that people are getting tired of AI slop, or the fear of missing out of having the latest AI tool or AI generated content that doesn't need to exist, and that actually, maybe people are looking for opportunities instead to re seek that human connection, which sadly, you know, is lacking in certain elements of society.
Yeah, I agree with that. Sounds to me that we should really take a student centred approach when we talk about critical AI literacy, and really have them at the centre of the conversation and hear from them and speak to them, but also understand the kind of impacts that tools like detection software can have on students motivations and how they think of themselves and how they see their future playing out.
If all these softwares are not working in their favour, or they're not working as as they're designed to.
So it sounds like we should really keep students at the centre of these conversations about critical AI literacy and about how these interventions should be designed, how we should lead the conversations around AI detection software, but also about AI literacy and how students should use these and when they shouldn't use them.
Thank you so much, Sam, for for being on the podcast with me today, and I really enjoyed hearing all your ideas and all your perspectives. And I have a lot to think about as well.
When I when I leave this podcast, I really hope that all our listeners will have the opportunity to read your article and to learn from it as well, as much as I learned from it.
So thank you very much for being here, Sam.
Thank you so much. And yeah, I just want to say this is not a plant at all, but I'm a huge fan of Hello World, and I read it almost religiously and recommend it to all my colleagues as an unbelievably useful source for how to think about, like computing and recently AI in the classroom as well.
So this was a genuine privilege to be on this podcast with you Rehana.
Thank you.
Thank you so much for tuning in today. If you'd like to read the full article that inspired today's episode, you can find it in the latest issue of the Hello World magazine.
Head to Hello World.CC to find the latest issues, but also to subscribe, explore our back issues and listen to more episodes.
That's it for today and we will see you next time.