Podcast: Lead us not into temptation: notes from the StudentXGenAI project – with Stephen Gow

This is a BALEAP TELSig podcast which I have had on my list to listen to for a while. The blurb: Last year HEPI reported 95% of students were using gen AI, but recent research from Stephen Gow and Sam Illingworth casts doubt on this figure. Today I’m joined by Stephen to talk through some of the key finding of his Leverhume Trust funded study that draws data from over 7,000 participants. What do students really think about gen AI in higher education, and how should this shape the way we treat it in the curriculum? As curriculum and materials development is a not insignificant part of my role, this is very relevant to my interests!

There’s a bit about Duolingo and streaks at the start but I’m waiting for the main content before I start making notes! If you are interested in Duolingo and streaks, listen to the start of the podcast! Though I do like the analogy between Duolingo/streaks and going to the gym every day but only using one piece of equipment.

It’s about what users actually think about the technology underneath the surface. Companies introducing no phone policies at work, offering lockers and pouches, to try and boost productivity. Not like picking up a newspaper back in the day because the sole purpose of the algorithm is to hold your attention. They mention teenagers wishing Tiktok could be uninvented. Technology use can be a compulsion. If you look/walk around in a library, you’ll see how many devices people have. You are in multiple places at once with it. When students used to come to university, the library would be the best information you could get and you’d have to go there to research things. The internet does give us access to far more information but managing that information is incredibly challenging. The prinicples of quality of information don’t change. The issue with GenAI as it’s come along is the “double-edged sword” that it does enable to a certain extent to digest all of the information but you aren’t too sure of the quality of the ingredients of that information. A model e.g. Claude may differ between different conversations/threads. Like the film Mickey 17 (maybe wrong number), he’s given himself up to be cloned for military and scientific experiment purposes. But each clone is slightly different. You could be having a really interesting interaction and then start up a new one which feels slightly different.

We have to keep in mind incentive structures: devices whose purpose is to hold attention and chatbots also are. E.g. if you finish a conversation, it will ask you if you want something else, suggest something you might want. There are models with guardrails on too. Tools will suggest ways to go to a student or researcher working on something. It takes a lot for you to say no. You can then discuss the ethics of doing the thing it suggested, and then it will say actually it will get you into trouble. But it can be useful if used in the right way. Just like with Duolingo. Ideally we want students who want to learn but even if the tools were to stop developing today, they have incredible capability and can do things that will massively affect peoples’ motivation to do things. The flip side is that the human brain is incredible, so how do we have a co-existence of those intelligences.

The elevator pitch of the study: The whole project is the StudentXGenAI project. The dynamic tensions paper is a scoping review, including trialling different AI tools. He approached it like a student would. Within the paper he talks about the findings from 40 papers limited to qualitative studies as the plan was to do interviews subsequently. Reflecting on student and researcher use of these tools. After the scoping review came 20 interviews with students. Then the Leverhulme funded bit was a survey of students in the UK. 7000 responses. They found a window without other surveys and got university buy-in by giving them the data of their student respondents to analyse. While for the project it was aggregated. They targeted students whether or not they use GenAI. 70% were using it, 30% were not using it for anything (studies, work, personal life).

Knowledge and access: do they know how to use it, how did they learn how to use it? Mostly trial and error. University resources are very low down. They are very useful if they use them but mostly they weren’t used.

What do you use it for? Brainstorming, computer coding, etc.

Attitudinal area of questions: Do you trust the outputs? What motivation do you have for using the tool (easier, faster, helps me to overcome blocks were common answers).

There was a petition at the University of Salford pushing back against the integration of GenAI.

There is a real tension between productivity/performance and learning. Instrumental action (Wnting to complete something to get it done) vs more communicative action where you are learning to do something and communicating with people in your society to get things done. If we are too instrumental, we stop communicating. We are in competiton with each other. Competition vs collaboration are another area of tension. Degrees in a massified system become less of a thing you do because you want to do it and more something you do because it is the done thing. A means to an end. So students more likely to approach it instrumentally. If you are stressed and getting into debt and find there is a tool that makes things easier, you will consider it a benefit.

The dataset was deliberately large in order to be able to draw some conclusions. The bulk of the questions were the same as an Australian study, and there were a lot of similarities which was interesting.

In terms of institutional policies, we have see versions of the traffic light system etc, assessment scales, two lane approaches, based on the assumption that everyone is using it so let’s facilitate that. Its efficacy as a learning tool is in question so there is a tension between creating conditions where we regulate it vs encouraging use of it because it’s the future.

Based on the data from the 7000 students, there is room for regulation. You can’t ban it but by having these systems you are saying it should be regulated. There is a false binary between using genAI and not using genAI. While the systems mentioned above, if implemented fully, at first assessements would primarily be red, and then over time there would be shifts into amber and green.

Staff don’t trust students, students don’t trust other students, staff don’t trust staff when it comes to using GenAI. The survey shows that students appreciate clarity and want to do the right thing. The majority of the students are honest and want to do the right thing. In both the Australian and this data, 67% of students will be honest and won’t use if when told not to. We need to hold onto this. The majority of people want to do the right thing. A significant minority will use it when not allowed. 10-15% self-report very instrumental approaches, just going to use it. This matches self-reporting on deliberate plagiarism and contract cheating. So what do we do, how do we find the balance between security of the assessments and trusting the majority of students. Security measures often impact the validity of the assessment. We need to take a step back and instead of having traffic light systems etc, have targeted experimental modules where staff and students work together and we see how staff and students are really using it. Have learning enhancement digital people working with staff and students, to see what really works. Then carry what is learnt to other departments. Increasingly there are wearables so we need to have a dicussion about privacy and data issues.

The students who are honest and recognise that using genAI isn’t the best approach, still have to deal with the instrumental pressure. They can see that other students are using AI and giving a false impression of ability and they don’t want to be at the tail end of the bell curve so they are pushed into using it rather than really learning. They don’t feel like they can afford to do the work in the way it was intended.

Is writing a prompt a skill in the same sense as the things that it is replacing? The Australians went back to pen and paper exams while in the UK we were still in a post-covid mindset of very, very enhanced trust – 6hr window, do the exam online, we trust you. If you look at pre-covid exam regulations and security, all you had was the person in the room and their brain. Post-covid, everything was more unsupervised, tools were not proctored. A worrying lack of attention to assessment security. What impact might that have on motivation?

Now it’s 2026, GenAI became a thing in 2022 so there has been time for the kneejerk outright ban, to embrace, and then trying to find a balance. The new normal. The challenge to trust of the digital world. The free era of AI is coming to an end now, as they need to finance, so time for premium. University of Manchester has given co-pilot to all students and staff, lots of universities doing similar. But more common for students to use free tools which have all sorts of data issues. Who owns the information? More collaboration between institutions and negotiating as a block with the GenAI companies is needed but we aren’t there yet. Everyone is trying to reinvent the wheel but it is not joined up enough.

What is the next big question that needs answering? To reflect on the perceived benefit of productivity. If a student can be productive and perform well without learning well, then there is a huge assessment validity issue. How do we solve this, and genuinely assess in the way that we learn?

Further reading

Chung, J., Henderson, M., Slade, C., Liang, Y., Pepperell, N., Corbin, T., Walton, J., Yu, AS., Bearman, M., Buckingham Shum, S., Fawns, T., McCluskey, T., McLean, J., Oberg, G., Seligmann, A., Shibani, A., Bakharia, A., Lim, LA., Matthews, KE. (2026). The use and usefulness of GenAI in higher education: Student experience and perspectives. Computers and Education Open, Available at: doi:  10.1016/j.caeo.2026.100347.

Gow S, Illingworth S (2026), “Dynamic tensions: an AI-assisted critical scoping review of university students’ qualitative experiences of GenAI”. Artificial Intelligence in Education, Vol. 2 No. 1 pp. 67–89, Available at: doi: 10.1108/AIIE-06-2025-0151 

Gow, S. and Illingworth, S. (2026) “It is a temptation to get it to do the work…” – student experiences of GenAI in UK universities. 09 Apr 2026. Advance HE. [Online]. Available at: https://www.advance-he.ac.uk/news-and-views/it-temptation-get-it-do-work-student-experiences-genai-uk-universities [Accessed 24th J 2026].

My thoughts:

It was a good podcast but I think I was hoping for more student voice to come through. More qualitative information from students relating to the three areas mentioned – knowledge and access, what do you use it for and the attitudinal aspect. I think I was also hoping for more concrete suggestions for curriculum and materials design. It may be that I didn’t successfully extract that information from the podcast format. I shall be reading the articles linked to above to glean more. I suppose that is kind of the point: the podcast is never going to be that comprehensive! (Edit: the “Dynamic tensions” one has some interesting stuff in the “Thematic analysis and discussion” part! As does the paper about the Australian study!)

Anyway, I think the bell-curve thing is very interesting – students using it even though they’d rather not because other students are using it and they don’t want to get left behind. That is a very strong argument for designing assessments that reward the skills you are teaching, accept ethical use of AI but don’t reward non-ethical use of AI. We have attempted to do that in our coursework essay assessment via reworking the criteria, with some success but the future will be impacted by evolution of the tools and further evolution of the criteria that is outside our control.

There also seems to be tension between AI-required assessment and the 30% who don’t want to use AI for various reasons (ethical, environmental, want to learn themselves). I suppose a carefully designed assessment could address some of that (e.g. make it so that they do still learn themselves; make it so that it doesn’t require heavy interaction with a tool so that the environmental impact is lower etc.)

It was heartening to hear that students appreciate clarity and the majority will try to follow the rules, as clarity and support is what we are trying to achieve in our curriculum and materials. Though I wonder if that is the case regardless of context. I.e. does the bell curve issue (or the need to succeed) have more influence in a context like ours where the outcome is very high stakes (determines whether or not the student gets to progress to univerity in the first place)!

I think it would be interesting to show students some questions and results (some are illustrated in the Advance-HE article) and discuss where they fit in/why, to get them thinking about it and give them insight into how others use GenAI. Perhaps get them to answer the questions on a google form then show the results side by side with the results from the 7000. All sorts of things you could do with it! Watch this space. 🙂

Trinitiy College London webinar: The Learning GPS – Helping Students Navigate, Monitor and Recalculate Their Learning Journey – Maria Eugenia Ianiro

Session plan:

  • Introductory questions
  • Introduce Learning GPS
  • Metacognition definition and importance
  • Activities
  • Final takeaways

Question 1: What do you before starting a journey? Imagine that you want to travel, what is the first thing you do when you want to visit a place that you have never visited before.

Participants’ answers: Look for a map, check the route, gather necessities, plan day by day, obtain information about the place, pack a passport, check for restaurants…

Maria suggests: choose destination, check the route, monitor the traffic, recalculate if necessary (if there is a problem/situation where we need to readjust).

Question 2: Do your students do the same when they learn?

Participants’ Answers: yes, no, not sure, not always, not really, it depends… varied answers!

Maria proposes: using a Learning GPS:

GGoal: Where am I going? So this is what we need our students to think about.

PProgress: Am I on the right route? Am I ok? Or do I have to make some changes?

SSuccess: How far have I got?

Maria says this framework should be accompanied with metacognition which is what we will look at now.

Metacognition = “thinking about thinking” (Flavell, 1979 p.906) It is the ability to think about and regulate one’s own learning. In the classroom it means helping learners become aware of how they learn, what strategies they use and how they can improve.

Why is it important? An essential element in today’s education but sometimes neglected by teachers as we don’t know how to implement it in lessons. It is an ongoing habit that transforms learning into understanding. Very important in age of genAI where there is a lot of content at hand. Metacognition CAN be taught, modelled and practiced and it SHOULD be.

G – Goal

Where am I going today as a learner? When students know the destination they can make better decisions, stay focused when they face challenges, and recognise success when they arrive. Every class needs a destination. We need to share that destination with students so they know where they are heading.

Question 1: How do you usually begin a lesson? (My answer: warmer, learning objectives; other answers: attention hooks, chat, emotional beginning, what do they already know about the topic…)

Question 2: Before your students start a task, do they know the activity…or the destination? (Varied answers!)

The activity and the destination are not the same thing. The activity is the what, the destination is the why/the goal.

Suggestions:

Here are some activities we can use to activate this idea of working with a goal.

My Learning Destination

Before starting, students complete:

  • Today I am learning to… (to is important.. so it isn’t just Today I am learning the present perfect)
  • This is useful because… they need to be able to connect what they are doing with how it will help them.
  • I will know I’m successful when…

Before I start checklist

Students tick before a task:

  • I understand the task
  • I know what the final product should look like
  • I know how much time I have
  • I know what resources I can use
  • I know what to do if i get stuck

This is practical for writing, projects, exam tasks or group work.

Success Criteria Builder

Instead of giving students success criteria, students build them. For example, before writing an email…

“What makes a good email?”

  • clear greeting
  • reason for writing
  • polite tone
  • closing phraase
  • correct punctuation

Students can then use this list to guide them while working to help them stay on track.

Choose your route

Give students different ways to approach a task.

E.g. learning new vocabulary

Students choose one:

  • make a mind map
  • group words by topic
  • write exmaple sentences
  • draw pictres
  • teach the words to a partner

Which route will you take today and why? (Activate metacognition)

Predict the roadblocks

Before the activity students complete:

  • I think this task will be easy because…
  • I think this task may be difficult because…
  • If I get stuck, I will…

Strategy passport

Before the lesson, students choose one strategy to travel with

Today I will use

  • underline key words
  • ask for clarification
  • use my notes
  • reread instructions
  • check a model
  • use a dictionary
  • rehearse bfore speaking
  • plan before writing

This kind of activity will help them become more aware of the strategies they use

KWL For Language learning

Before a topic:

  • K – What do I know?
  • W – What do I want/need to know?
  • L – What did I learn? (complete at end of class)

Plan my answer

Before speaking or writing, students complete:

  • My main idea:
  • Two words/phrases I want to use:
  • One grammar point I want to practise:
  • One thing I need to be carefrul with:

Traffic forecast

Before the task, students predict their confidence:

  • I don’t know how to start yet (red)
  • I may need some help (amber)
  • I think I can do this (green)

What would help you move from red to yellow? yellow to green? Then they need to think about their strategies

3 minute planning pause

Before any main task:

  • Minute 1: What do I have to do?
  • Minute 2: What do I already know that can help me?
  • Minute 3: What strategy will I use first?

So they need to think abotu the activity itself and what strategies they will use to complete it.

Goals: Students don’t need to know only what they’re doing. They need to know where they are going.

P – Progress

Question we need to ask: Am I on the right route? Why is this important? Learning doesn’t happen in a straight line. Effective learners constantly monitor their understanding, notice when something isn’t working and adjust their approach. If we help our students to do this, we will reduce the level/amount of frustration students have. It is ok to make mistakes and for things to go wrong but we need to be able to adjust or it is really frustrating.

It isn’t about asking “am I right?”. It is asking “is this strategy working?” “Do I need to make a change?”

Ideas:

Learning dashboard.

Instead of checking only answers, students monitor four aspects of their learning.

Self evaluation:

I am….

  • understanding
  • participating
  • using strategies
  • staying focused.

For each one, students choose where they are on a scale of smiley, neutral, sad. Then after they do this, ask “Which area needs your attention?”

GPS recalculating

Whenever students feel stuck, they cannot simply say:

“I don’t know”

Instead they complete:

  • I am stuck because…
  • The strategy I’ve tried is…
  • Now I’m going to….

E.g. I’ve already reread the text, now I’m going to underline the key words.

Midway Pit Stop

Halfway through the activity, students answer 3 questions:

  • What is going well?
  • What is still difficult?
  • What will I change during the second half?

One minute. Then continue working.

Then at the end ask: which area needs your attention?

Strategy tracker:

During the lesson they tick every strategy the actually used.

  • reread
  • asked a question
  • used a model
  • looked for examples
  • worked with a parnter
  • checked my mistakes
  • slowed down
  • explained my thinking

Which strategy helped you the most?

Strategy checkpoint

Did my strategy help me reach my destination?

  • How did planning your ideas help you write a better text?
  • Why is it important to link your ideas in and between paragraphs clearly?

Students have to think about the strategies implemented and check if the strategies they used help reach the destination

Confidence thermometer

Every 15 minutes students indicate:

  • I’m getting lost
  • Getting there
  • Very confident

Then ask what would move you one level higher?

Learning selfie

Pause.

Students finish the sentences:

  • Right now I feel… because… .
  • What I need next is…..

The selfie is an invitation to post and reflect on this particular moment.

My next best step

Instead of what should the teacher do? Students ask “What should I do next?”

Possible answers:

  • reread
  • ask a classmate
  • check my notes
  • practice another example
  • ask for feedback
  • keep going

Error detective

While working, students identify one mistake. Instead of correcting immediately, they ask: Why did I make this mistake? What strategy could prevent it? This makes mistakes part of the learning process positively.

Checkpoint cards

Every 15 minutes, students answer one question:

  1. CHeckpoint 1: What have I understood?
  2. Checkpoint 2: What strategy am I using?
  3. What should I keep doing?
  4. What should I change?

To help them be aware of what stratgies they are using while completing activities.

Stop – think – adjust: a simple routine that can become automatic in your lessons.

  • Stop: students think “What am I doing?”
  • Think: “Is it working?
  • Adjust: “What should I change?

Mid-journey questions – these help students do the same: pause, think and decide what to do next.

E.g. How does the way we read a text change when we use information from it in a new way? How can you practice this skill outside the classroom?

The three mirrors:

Students monitor learning through three lenses

  • Content: do I understand?
  • Strategy: is my strategy helping?
  • Effort: Am i giving my best attention to this?

S – Success

Have I reached my destination? What have I learned about my leanring? Succes is not completing a task, it is understanding the journey that led to it. So the most important thing is the learning process of the student.

The rear-view mirror:

Project an image of a rear-view mirror. Students complete:

  • Today I learned…
  • I was surprised by…
  • The most useful strategy I used was…
  • Next time I’ll…

This table can be related to strategies or techniques implemented in lessons.

Before and after

  • At the beginning of the lesson: I think…
  • At the end: Now I think…

Learning highlight

Students choose one moment. Concentrate on that moment and complete the sentence: Today the moment I learned the most was…because… The most important part is the wyy. This is encourages them to identify high impact learning experiences.

Strategy awards

Give students categories:

  • best strategy
  • biggest improvement
  • most helpful mistake
  • best question
  • most creative solution

This moves the focus away from marks, to focus on different areas to evaluate.

My future self:

Finish the lesson by writing:

Dear Future me,

Next time remember….

This creates transfer, which is the highest level of metacognition. Instead of only reflecting on the past, students prepare themselves for future learning.

A postcard to my future self:

Today I discovered… I hope you remember… See you next lesson!

Exit Ticket 2.0

Instead of Today I learned…

Use:

  • One thing I can do now that I couldn’t do before…
  • One strategy I’ll use again is…
  • One thing I still need to practise is…

The next destination:

Finish every lesson with:

  • Today I arrived at…
  • My next destination is…

Ask students to set their own goals. This reinforces that learning never ends. With these two sentences, you are adding metacognition.

Reflective hexagon

Students write six words:

  • What I learned
  • What helped
  • What challenged me
  • What surprised me
  • What I’ll remember
  • What’s next

My learning story:

Students complete:

  • At first…
  • Then…
  • Finally…
  • Now I know…

Metacognition promotes deeper learning, enhances self-regulation and motivation, fosters lifelong learning and supports 21st century skills.

My thoughts:

This was a great, if exhausting, session! Loads of great ideas that are really simple, I hope to try some of them out next academic year, in individual lessons and also in conjunction with the end of week reflection form I do with students. As ever, just need time to think about the hows and whens of it all!! 🙂

ELTC TD Session: Reflections from BALEAP ‘Gen AI PIM’ Event – Tim Radnor

One of our teachers attended the recent BALEAP PIM: EAP and the Academy in the Age of GenAI: Implications for practices and practitioners and today is sharing his reflections on it. It will be interesting to hear about this and compare it with the AI-related talks that took place at the InForm conference that I attended sometime after this PIM took place!

The session will include a few reflections based on current research: focus on 2 conference presentations that Tim attended, on two aspects of how Gen AI and EAP intersect. We will also discuss how the issues raised may impact our professional practice.

We started by looking at statements from a session Tim gave in November 2024 and identified how things have changed since then:

  • There were opinions that AI is even more widespread now and established now, and getting better at what it does, with more natural responses.
  • There is also more opportunities for creating AI tools e.g. Gemini Gems, so more customisable.
  • In terms of policy, it is on the way to become embedded in assessments and curricula here at Sheffield, but across universities policies are very different, there isn’t uniformity.
  • We are more aware of drawbacks including environmental damage, ethical issues, reliablity of results, information ‘dumping’.
  • Academic research into Gen AI and EAP has grown rapidly in the last two years, is much broader now.

In terms of the BALEAP day, there was a full day of presentations with 8 strands that each included several papers. There was a lot of overlap between strands. We are going to focus on two only:

First presentation: How Gen AI reshapes perceived academic writing ability in EAP contexts by Dr Elhadj Moussa BenMoussa (University of East London)

Tim had seen one of Moussa’s presentations before in the academic literacies BALEAP sig so knew of his research. Moussa’s study looked at student drafts of academic essays, then AI-assisted revisions of the same texts, as well as tutor feedback on draft and edited version. (Tim doesn’t know the level of the students or class they were in.) But they were in-sessional. And recurring patterns were found:

Before AI: writing weaknesses were visible and tutors could identify various issues e.g. grammatical errors, spelling, cohesion.

After AI: the text improved, weaknesses such as those above became hiddent beneath fluent prose but the thinking had not improved.

“The illusion of competence” was the speaker’s phrase. AI texts often appear fluent, coherent, academically structured (in terms of paragraphing) and confident in terms of style and tone. However, they may contain weaknesses in terms of argumentation, critical engagement, source integration, disciplinary reasoning and authorial stance (writer’s voice in other words). The speaker was looking at it from an academic literacies perspective. He found that the text appears stronger but the underlying understanding remains unchanged.

The AI version has added ideas, has better cohesion, higher level vocab. In the AI-assisted version, the element of the student’s perspective on feedback has been lost. If anything, the student draft is more critical even though simpler. The AI version has no argumentation in it. We are assuming an average level of simple, straightforward prompting. Prompting expertise doesn’t contradict the argument here however. The AI assisted version sounds more academic but does it demonstrate deeper knowledge? No, it is just descriptive.

Why does this matter to EAP?

Traditional indicators of proficiency are becoming less reliable for assessment as AI performs these automatically. EAP practitioners are uniquely places to help understand the shift to evaluating reasoning, knowledge, rhetorical decision making, source use and intellectual ownership.

Moussa argues that we need to “make thinking visible” (epistemic visibility was the fancy term used!). We need to make students demonstrate why they made choices, how arguments were developed, how sources were selected, how evidence was interpreted and how conclusions were reached. I.e. shift from what was written to how was knowledge constructed.

As a takeaway, Moussa recommended:

That was the first presentation. The next stage was to discuss for five minutes: what do you think the implications of these ideas are for how we adapt assessments, feedback on written work and classroom activities on EAP (pre-sessional, in-sessional, foundation…)?

Ideas shared in our session:

  • focusing on speaking rather than writing as more difficult to fake
  • focus on the process rather than the end product but at the same time do we need to go back to in-room exams rather than ongoing/project assessment? Not sure how much focus on process helps.
  • Have to accept they are using it and think more about how
  • More use of reflection on own learning progress with reference to specific lessons and learning materials
  • Given that students will be expected to be able to use AI effectively in future workplaces, it arguably makes sense to assess their ability to use it well rather than how they perform without it
  • A variety of assessment types to cross-reference for gauging student abilities
  • One of the things is in class helping to develop their analytical skills on the spot e.g. when we do a task in a breakout room, saying to students right now you’ve done that task, how easy or difficult is it (scale of 1 to 5), how do you rate yourself? (1 to 5). Make them think about it. Then, so why was it difficult? Why was it easy? To get them used to providing answers that aren’t AI-based, but are their own analysis of their learning.
  • Also for feedback on regular tasks, you can build in that critical process. Moving beyond lanaguge correction to looking at the ideas themselves and encouraging reflection.
  • Consciously use AI in a course programme and ask students to use it in a particular way and then they critique the output that they get
  • Adapting marking criteria
  • For each assessment it has to be stated what the parameters for AI are.
  • Bit of a paradox: that shift to arguing, styles, evidence etc sounds as if we have surrendered and are trying to engage students receptively but not productively?
  • Or is it more helping the students improve their reasoning etc as they produce the text and then polishing it with AI? If you can help students as much with reasoning and argument as with language, but you have to look at language as well e.g. how to do these things. But if you already have the reasoning formulated in depth, the polishing is less of an issue. The suggestion isn’t to copy and paste!

Focus on Gen AI literacy.

What do we understand by this term?

Ideas shared in our session:

  • writing prompts
  • being able to evaluate output quality and relevance
  • being able to use AI tools to enhance learning and productivity (ethically and appropriately)
  • knowledge and flexible application of AI tools for different purposes, i.e. social , academic, vocational activities
  • know how to use accurately and appropriately according to circumstances

Here is a definition of GenAI literacy from the literature:

Second presentation: “Student practices and experiences of Gen AI: developing a practice-oriented heuristic for GenAI literacy” by Julio Jiminez, Katherine Mansfield and Richard Paterson from University of Westminster.

Background: There is tension between institutional discourse (polices, integrity, guidance, how we teach it, recommendations) and student practice. Students are already integrating gen AI into their practice but in more nuanced ways than institutional discourse assumes. Policies struggle to keep pace with rapidly changing practices and technologies.

Research question: What do students actually do with GenAI in everyday academic life?

The speakers surveyed 441 students across 66 disciplines.

Findings:

GenAI is used by students for…

  • generating ideas (68%)
  • planning work (66%)
  • finding sources (33%)
  • receiving feedback (24%)

Note: Very different percentages between the top two and bottom two. They talked in focus groups with the students as well.

Key themes:

  1. Students use GenAI strategically: they pick specific tools for specific tasks; their selection takes into account what they consider the uses and limitations of the tool; their use suggests that rather than simply using tools because they are there, use is based on the fit between the task and the tool.
  2. GenAI literacy is emerging: students do critically evaluate the accuracy, reliablity and bias of tools and output and they are developing awareness of what is and isn’t appropriate use. To improve this, there needs to be ongoing decision making which is context-sensitive.
  3. Learning is increasingly human-AI mediated: GenAI is used to support planning, drafting and revision. However, students maintain their evaluative control over the outputs. Therefore, academic practice is becoming a human-AI hybrid process.

Overall, this all indicates that students’ GenAI literacy is moving beyond technical proficiency towards strategic, critical and context-sensitive use.

The speakers came up with 4 principles for how these findings could feed into approaches to teaching GenAI literacy:

They came up with this framework (Framework rather than practical steps):

From left to right it moves from more basic to more complex.

Finally, we discussed the following more practical questions based on this session:

  1. Does any of your experience of student use of GenAI align with or differ from the findings in any way? How?
  2. How well does the heuristic (framework) match how students are taught AI lteracy in an institution or teahcing situation you are familiar with?
  3. Can you think of any practical exercises, workshops or courses for sudents that could incorporate the ideas from this session?

Ideas shared:

  • Students are given specific information about AI use but it isn’t incorporated into day to day activities. So one participant uses it in classes to help students become more independent users but class time is limited. Need longer sessions where you can demonstrate something and give students time to try it. So it shouldn’t be an add-on anymore.
  • Students used gen ai to create diagrams for their presentations and cited them as well as referenced their creations They were relevant and much better to support their speaking than looking for what approximates relevance for their topic – So creating visuals. It worked really well in conjunction with academic referencing/citing. They were responsible enough to cite and add sources if necessary. It helped them support their speaking. This requires increased literacy in terms of knowing about doing that and how to do it. However, with simple prompts they managed very good results.
  • Get Gen AI to write an essay in class and then get the students to analyse its weaknesses, such as poor argumentation and use of evidence. This requires criticality in this area.
  • Often a massive disconnect between AI-positive message from here and what people experience. There needs to be a more joined up approach.

Tim’s overall reflections on the conference were as follows:

  • he doesn’t know as much as he thought about EAP and AI (I think that is probably true for many of us!)
  • There is a lot of research going on into EAP and AI
  • EAP professionals have an important role to play in the transition to the integration of Gen AI in HE institutions
  • Students may be using GenAI in a more strategic, critical way than EAP practitioners assume

My thoughts:

My first take-away is how much student agency there is around AI. That whole strategic use thing. It suggests that it is really important in terms of AI literacy and how we teach use of AI to involve the students’ perspective as much as possible, to better understand their use and be better able to help them shape their use to fit the specific context in terms of what is and isn’t acceptable and helping them refine their use. I suppose the more confident we become about GenAI, the more confident and better able we will be to have these sorts of conversations with students.

Thinking about my specific context, this last academic year has been the first year where we have acknowledged AI rather than ignoring it beyond having a blanket ban which is impossible to enforce. We relied mainly on interactive content put together by the digital team, which we gave students time to complete in class. In the coming year, in contrast, the idea is to have in-class activities that are teacher-led rather than interactive content-led regularly in the first five weeks of the course and then at a few specific points thereafter. We will see how it goes!

My second main take-away is the difference between the student example and the AI revision. I can imagine a student who fed in the studente example and got back the AI example would be like “yeah, great, this is better! I’ll use this!” while their actual message has been lost. Of course, they’d need a clear understanding of both their message and the AI output in order to realise that. So I wonder how we help them become better able to do that. I need more time to think about that! I always need more time to think about things but time is in such short supply…

Gen AI: Gemini Gems and Notebook LM

Another university-wide training session relating to AI, which took place on Friday 22nd May!

Gemini Gems

Gemini Gems can be used to create custom versions of Google Gemini. They are a custom AI assistant. Gems are able to provide a more tailored and relevant response than a standard Gemini chat. It is a way of training Gemini to produce output that is more tailored to your needs. It does take time to create them properly but it could be worth the time required for using/making if you are always giving the same instructions to Gemini. E.g. Students could use it to make a personalised study assistant. Gemini also has some ready-made ones that you can edit. To share a Gem, you do it in the same way as other Google products e.g. docs/sheets (but you have to enable” Smart features” to do so). Ensure you make it view only.

There are limitations: Gems draws from the whole internet. You can feed it information – knowledge – to draw on, but if you ask questions outside that remit, it will bring in outside information just like Gemini does. And just like Gemini, it may hallucinate. There are ethical concerns about power consumption, copyright and data protection. Don’t put any sensitive data into it. While at the moment the university version is closed, who knows what will happen in the future, how that information could be used. Where it asks you to put in information, it can only be a pdf or similar file. Not a webseite. It is very easy to create something that almost works but a lot harder to get it to work properly.

Of course, we do still need to critically evaluate the output. If you don’t use Gemini that much, you might not find Gems useful as it might not be worth the time input required to set up effectively.

Notebook LM:

Notebook LM is a virtual research assistant. Instead of being based on a large language model, it is a retrieval-augmented generation (RAG) tool. All responses are based on the sources you provide, rather than drawing on everything available online. It is a double-edged sword – it is really useful for looking at large quantities of text etc but it could quite easily do your work for you. So it is important to develop effective use of it.

You can upload sources, interact with them via a chat tool. It will ask you to upload your files/content. Click insert. You will see the documents appear that you upload. Without being asked, it will give you a summary of what is in the documents. Those documents are your knowledge base. Then, you can give it an instruction e.g. “I am an educator working in UKHE specialising in [digital education]. Draw out five key points from the sources relevant to me and present them in succinct bullet points.” and it will carry that out.

It is a little language model of your own that you can interrogate as you wish. E.g. find three areas of disagreement between these sources. So if you have a long document to explore, you can ask it anything you like about it. You can untick a source if you no longer wish it to be used in the responses. It will always keep trying to please you, ask you questions, engage you but it doesn’t hallucinate content that isn’t in the sources you have uploaded.

You can also do this: “I’m an undergraduate student writing an essay on ‘Identify the pros and cons of AI use in higher education’ – can you write me a 500 word literature review?” – obviously a danger here that you are offloading your work onto the AI and then if you use that information, it is ethically questionable in terms of false authorship. In “studio” you can ask for other outputs e.g. an audio file/pod cast. These can take a while to generate. You can get a 1 – 1.5 mins introductory piece or a 10 minute discussion between two voices. Students could put notes they have made into LM and get it to make audio out of it to help with revision. This could be useful for neurodivergent students. It can also turn it into a slide-based summary. All of this is entirely based on the texts you have inputted. A slide-based summary could be useful for a visual learner. What you can’t do is edit the output. It is what it is. You/the student can also generate quiz cards or mindmaps to help with revision. It gives you more options of ways to engage with the material other than reading huge bodies of text.

In terms of limitations: you need have knowledge of the information you upload in order to evaluate/trust the output. And just like other AI, there are the usual ethical concerns, copyright and data protection issues. You should avoid using it with unpublished draft papers or other confidential research materials, as well as any sensitive information.

Love it or hate, it is available to all students and staff here! It does things that we don’t want students to do but it does save time and we can’t stop people from using it. So, we need to try and help them engage critically. Hopefully our increased awareness helps us be better able to do this. Use with caution is probably the best description.

BALEAP TAFSIG Webinar ‘Developing a viva-style assessment for a large language course: a collaboration between two large scale pre-sessional programmes’ –

I was keen to join this as in title it seems to be relevant to our ACP – Academic Coursework Presentation – which has a Q&A component. TAFSIG is Testing, Assessment and Feedback Special Interest Group within BALEAP (glad she clarified that because I didn’t know!) They have a YouTube channel with weinar recordings (I shall have a look at that!)

4 speakers: Craig Davis, Nicola Harding (Manchester), Lori-ann Miln, Phillippa Bunch (Southhamtpon).

Context and collaborative development:

The presessionals are similar in some respects: large student and tutor cohorts (100+ tutors). Both assess writing in an essay (challenges brought by genAI and suspected malpractice but no real way of authenticating engagement with writing components).

Manchester: Reading to writing (80% writing, 100% reading) L2S (50%S seminar, 100% listening). Previously seminar with brief presentation followed by discussion was 100% of speaking score.

Southhampton: Assessments are more separated out. One key difference from Manchester is that students can choose their own topic for their researched essay = a lot of different research areas. Reading is demonstrated through writing. Speaking is a presentation followed by Q&A. Listening is a lecture followed by discussion where skills are assessed. There are different read/writing tutors and speaking/listening tutors.

Manchester and Soutthampton had a collaboration, identifying similar challenges a few years ago and sharing ideas around assessments and after several meetings realised they were all moving in the same direction, as GenAI use became more prominent and they were trying to deal with that in the assessment design. They were also prompted by feedback from lecturers across the university, saying that rehearsed presentation was not effectively assessing their ability to produce spontaneous speech.

Manchester Case Study

Created the question and answer assessment:

  • Question 1: Prepared question, same for all students (2 mins)
  • Question 2: product focused from question bank (3 mins)
  • Question 3: process focused question seleted from question bank (3 mins).

All students were given the same question: To what extent should AI be integrated into HE?

Example of Product focus: Tell us about a specific source you used in your essay – how and why did you use it?

Example of Process focus: What was the most significant/helpful piece of fb you received and how do you respond to it? (make sure student answers both parts of the question)

Follow up prompts: Why do you think that? /Is there anything else that you found helpful/unhelpful.

Scaffolded process:

  • Stage 1: Assessment overview (brief introduction during student orientation talks)
  • Stage 2: Tutor-led sesion focusing on the assessment (every assessment is supported by synchronous taught session, look at criteria and apply them to example answers).
  • Stage 3: individual tutorials x 2 redesigned to follow a Q&A format (existing activity, part of the course, mini QAA for that to enable practice.)
  • Stage 4: individual study (checklist of preparations and a reflection leading into final stage)
  • Stage 5 (final questions and preparations: final group tutorial, centres on the Q&A and the assessment coordinator drop-in – where students could ask any questions, not many did)

How were tutors supported?

  • The assessment was piloted with volunteer students from PS April 2025 and standardisation packs were created using videos from the pilot.
  • An answer guide (University of Manchester is quite prescriptive around this part of the assessment: pros and cons, but one of the pros is the easy ability to produce such a guide!) was created to support marking.
  • A streamlined criteria for live marking: was made as useful as possible for the tutors, as there are lots of students marking up to 18 students each, and doublemarking.
  • Marking template was provided with space to write which questions were asked and space for notes about the answers.

Students were assessed on speaking by focusing on: fluency, pronounciation, language. Students were also assessed response to the question by focusing on: relevance, specificity, development

Speaking (fluency, pronunciation, language) contributed only to speaking (25%) and Response to questions 1-3 each contibuted 25%. For writing score, 20% writing and then response to questions 1-3 contributed 33% each. So increasing the focus on process.

They were quite prescriptive in terms of what was provided to tutors, as it was the first time to run it large scale. Everything was live doublemarked, assessments all scheduled over 3 days.

How did students perform?

Generally fairly consistent. Outliers: Speaking: 57 students scored 20%+ higher on L2S than on Q&A. 12 students scored 50% or below in QAA but above 70% in L2S Writing: 20 students scored +25% lower on the Q&A

Overall, the assessment was well received by teachers and students. However, timetabling was a big challenging as everything was live doublemarked. Questions banks supported tutors and ensured consistency, while prompts supported students in speaking for the full 3 minutes (mostly). However more question-specific follow up prompts and more guidance on managing the discussion elements could be needed.

Part 1 and 2 had a lot of overlap and also some scripting and reading from the essay = more difficult to authenticate and mark, as not saying a lot about student engagement. Question 3 was more revealing about how students engaged with the process, and therefore more useful.

Next time: They want to switch the focus to process rather than product – product – process. Still the same components (1 rehearsed, 2 not). This will require revisiting the question and prompts. They may also change the ratings to make Q&A 40% of writing. In terms of identifying outliers earlier = having a minimum component threshold for referral, so if student falls below 40% for any component they will be flagged. They also want to have a gap cap between the writing and the speaking

Southhampton Case Study

Soutthampton did a formative presentation (4 minutes) and Q&A (2 minutes) then later a summative presentation (6 minutes) and Q&A (4 minutes). The Q&A consisted of 2 questions: 1) demonstrate understanding of content and 2) demonstrate reflection on research process.

They don’t have much quantitative data at the moment but plenty of qualitative. The Q&A was worth 20% of the assessment criteria with content, structure, communicaton, precision and accuracy also each worth 20%. Students needed to be able to talk about they used their sources and how they used feedback.

In terms of tutor support: they provided structured tutor training embedded in the induction programmes. They also provided a question bank to support consistent assessment delivery. Finally, they did standardisation sessions. They thought this would suffice. But, after moderation and observation of formative assessment, they saw a lot of variability and disparity/inconsistency in tutors’ ability to run the Q&A in terms of formulating suitable open ended questions, and scaffolding student responses in real time. There were also struggles around sustaining interaction beyond surface level clarifcation, in terms of not allowing enough time and space for students to develop detailed responses. Some interactions were very brief, while others were more developed and encouraged critical thinking, thereby resulting in a better score. So, based on the formatives, in preparation for the summatives, they produced an enhanced guidance and question bank with initial questions and possible follow up questions, which helped with the issues identified when it came to the summatives. It should be noted that the questions had to be able to fit everyone’s essay topc even though they would all be different.

Things they found interesting: students developed more confidence in oral academic English, there was stronger evidence of research engagement, fewer formal academic misconduct cases but also the quality of student enagement depended heavily on the tutor’s quetsioning technique. A positive outcome: students reported feeling more prepared to communicate in an academic environment.

Looking ahead, they want to perhaps shrink the prepared presentation and extend the Q&A, with increased emphasis on process, linked to student folder. They also want to enhance tutor development with targeted training.

Key considerations and shared findings

  • Both assessment designs responded to AI by increasing emphasis on more spontaneous and more authentic Q&A
  • Different approaches but common challenges particularly around questioning, interaction and consistency.
  • Overall positive student outcomes.

Ongoing debate on balancing the structure for fairness and reliability with flexibility for authenticity and responsiveness.

Discussion/Questions

The first question was about student grading: how much focused on subject knowledge, language etc. Manchester doesn’t score much for content, focuses more on language, and relevant response to questions (which is sort of content based).

The second question was about the degree of mitigation of AI use. The focus on process has helped, moving away from end-loaded assessment and building it in throughout the course as well as building in meaningful dialogue from day 1. Also makes feedback into more of a process as it is revisited. At each formative stage there is opportunity for discussion with the students and this can be very constructive. Also, knowing from day 1 that they will need to do this encourages engagement with the process. With online courses, the live transcription thing is a challenge. Live Q&As are much better quality than online, as students have to be more natural and spontaneous and use oral strategies when they aren’t sure what to say etc. Any suggestions to help with online are welcome!

Reading/writing team do a lot of critical reflection as part of the course. (This related to a question about how prepared students were for the level of criticality required by the speaking assessment.)

The question bank: Soutthampton are asking 6 questions each time formatively in reflective tutorials so students get used to the style of questioning. A list of questions is not given in advance, but they do get two practices so students get the chance to practice responding to that style of questioning process.

How important is the prepared part of the presentation, would it be better to get more quickly to the Q&A. Answer: good question, they have been tempted to do away with the presentation part but they keep coming back to the point that on most degree programmes ss have to give a presentation, whether or not they use AI to do that, so the presentation skills are stilll useful to take forward and therefore it still has value. Focus on process is very present already in the writing, this may be brought more into the speaking as well. In terms of keeping the presentation, the prepared part allows more confidence coming into the more spontaneous part. There is still thought about changing the weighting so that the prepared part has less weighting.

My thoughts

Phew! Interesting hour, well spent! Southhampton’s current approach seems more similar to ours – except we have 7-8 mins presentation and 2-3 mins of questions. We have a mock and final, which I guess equates to the formative-summative. The latest cycle did identify issues like those mentioned above around consistency in questioning, in terms of difficulty of questions asked and depth/length of interaction. We had already discussed the need to standardise this more, so there are some good ideas in this session to draw on!

I feel that the Q&A component of our criteria could also use fine-tuning, drawing on some of the ideas shared today. At the moment the Q&A is only worth half a criteria but that is not something we can change without a Studygroup-wide discussion and process so it is definitely not a very near future thing, but we definitely can improve on how we do and mark the Q&A part, then perhaps at some point if we can shift it towards representing more of the overall presentation score (which itself is 50% of students’ speaking score), we will be in a better position to do so.

For now, we have an entirely different kettle of fish spilling all over the place development-wise, however! 😉

Assessment and generative AI

This was an internal workshop which was aimed the whole university not only us ELTC folk. Being the end of marking week and my work week (final hour thereof!) it was a good chance to grab the opportunity to do some development (woohoo!)!

The plan was to reflect on the impact of gen AI on assessment, hear about the university’s new common approach to gen AI in the curriculum and then focus in on what fair and appropriate use of gen AI in the context of an assessment might look like and what an “AI-required” assessment might look like.

Here are my notes from the session:

Currently, students can gain a passing grade or even a high pass, using AI in a written assessment. This use is not something that we can detect accurately and fairly, and therefore we can’t ban it as the ban would not be enforceable. However, that does give us the responsibility of ensuring that students use AI well and effectively, and develop appropriate knowledge/skills. We also have to prepare students for workplaces where AI will be used (though as yet what exactly this looks like is unclear!). So, rather than focusing on the negative impact on assessment, we should use these developments as an opportunity to focus and reflect on what, how and why we assess, with a view to improving the process for all concerned.

By Sept 2028, all programmes at the university will include 2 summative AI-Required assessments annually. Well, not *all* – this applies to Year 1 and 2 of undergraduate programmes. “AI free” and “AI-required” assessments are to be explicitly categorised. In AI-free assessments = it should not be possible to use AI, and skills must be demonstrated in an environment where AI assistance is not possible. The choice to do this type of assessment should be grounded in the learning outcomes e.g. communication and interpersonal skill evaluation, development of oral defence and articulation skills, verification of minimum competence, preparation for high-stakes professional contexts, simulation of professional practice conditions. In AI – required assessments, students would have to use AI in some way in order to complete the assessment and develop AI literacy in the context of their subject.

For all assessments that are not labelled either “AI-free” or “AI-required”, fair and appropriate use of AI should be possible (as long as it doesn’t constitute “false [automated] authorship”) in a variety of ways. We should include a statement on AI and academic misconduct (see below) on all assessment briefs. Alongside this, we need to teach students what good use of AI is in the context of a given assessment. However, this should not just be a list of what is and isn’t approved, as that is not something we can enforce in terms of the things on the “what isn’t” list.

There is still quite a lot of grey area – how to define “mostly” or “entirely” for example – where is the line? It does require unpacking for us to understand it fully and communicate to students effectively what it means.

Next, we looked at a generic task and what would be fair/unfair use of AI. Using AI to support elements of the task that the student is doing themselves and getting feedback of what they have done, are examples of appropriate use. A question that arose from a participant: Does getting AI involved in this way minimise student development in terms of the editorial process? Or, does it help them develop the ability by showing them a model of how to do it? It was suggested that perhaps it depends on whether the student lets the AI have the last word. That is, if the student simply adopts the AI comments wholesale without critique, then possibly the overall effect on the development of editorial process is negative but if the student approaches it critically, and learns from it, perhaps it can be helpful?

How could we communicate effective use to students? How could we adapt the task to discourage inappropriate use? Need to be careful of “traffic light” systems because they are not so suitable as we can’t actually enforce the “red” area rules. We need to talk to students about it, come to a shared agreement with them. There is also an issue that students have access to widly different AIs, of varying power. For AI-required assessments here, Gemini should be used – but again, how do we enforce that? Answers on a postcard! AI isn’t limited to LLMs, a participant suggested we need to consider what other platforms/programmes that exist and might support students. The response was that in large part, the focus on LLMs is because they are applicable to everybody, while more specific AI have more specific subject applications. To consider others, you need that more specific knowledge and skill set. (Which is an issue for departments and department-specific training I suppose!!)

Currently, there are some things that students can do better than AI but how long will it be until AI does do everything well in relation to the task? Then we won’t have that option anymore. Probably we are not that far from it. Is it worth putting in for major changes to an assessment when it takes 2 years for that to go through by which time the changes may be obsolete? E.g. hallucinations are getting rarer and it is less likely for it to use invented sources. A more current problem is it won’t go for the best journals or do a great job of assessing what is good or not. But then, you can get round this by specifying in a prompt what you want to be included/excluded. Assessing students in real time, using pen and paper, rather than digitally is an option, as is making the task more personal, having to draw on student experience rather than a generic case.

What about AI-required assessments? AI use might just be a component of an assessment, not a whole assessment that is about use of AI. It doesn’t even have to include student use of AI tools. They could look at pre-created content, talk about why they have decided NOT to use AI, analyse use of AI in a specific context, create a plan for how AI could be used ethically etc Students can ethically object to using AI but should still be able to learn about it.

So what might an AI-required assessment look like? What we need to start from is what skill do we want them to develop by using AI in this assessment? The university are putting together an AI literacy bank, broken down into Awareness, Competence and Ethics, which we can use to help inform task design/adaptation. In our context, students are in the very early stages of their academic journey so we would need to pick out what these students would most benefit from at this stage in their academic journey. [This is useful information as we were wondering if and how this would apply to us. Looking at the framework, it should be possible to identify which element(s) are possible for us to integrate and assess, building on what we have done already in this direction.]

We were asked to suggest modifications to the example task, and responses were more muted – possibly because it was quite complicated and there wasn’t much time!

It was noted that an AI-free assessment should be AI-free for pedagogical reasons, we should have a pedagogical justification. It is a descriptor rather than a label to adopt because we are scared of AI and student use of AI. As we start adapting assessments, need to think about what sort of changes – administrative, minor or major changes – are required. The reality is, the majority of changes relating to AI would hopefully fall under administrative changes: changing how the task is written, changing the information in the assessment brief provided to students = relatively quick to make. Changing an assessment type completely would then be a minor adjustment and require a longer process. A major change is to the programme level learning outcomes, so is unlikely to apply here. So anything around assessments would be classed as at most a minor change. We’ve got two years lead time – January 2028 ready for September 2028 intake would be when to put through minor changes to assessments. [Obviously for us on the AES (Academic English Skills) programme, it is a bit different as Studygroup timelines and processes are involved!]

It was an interesting session. Currently, our coursework assessments are neither AI-free or AI-required, and we have tried to encourage students to use it ethically to support their learning in the context of a given assessment rather than to do it for them. We have adapted our criteria for the writing coursework (extended essay) so that students need to do the skills we teach (critical evaluation, synthesis etc) well in order to score well, and conversely can submit a polished piece of work grammar/vocabulary-wise yet still score poorly. As mentioned above, how long the new criteria will hold effective for is anybody’s guess.

Based on today, and on the trajectory of travel for our coursework speaking assessment (presentation), which is moving towards greater emphasis (in terms of time and marks awarded) on the q&a part than there is currently, I wonder if that is where more personalisation could be incorporated and/or where we could get students to elaborate on how they have used AI, or any decisions they made around their AI use? Anyway, we shall see. In terms of how we teach students how to use AI effectively, going beyond checklists, that is one of our development aims for the coming semester – integrating that teaching more fully into lessons rather than it being more of a bolt-on of do’s and don’ts.

There will be more of these kind of sessions in the nearish future, so I will be interested to attend them!

“Let me hear the real you” M.E.T. webinar by Mark Heffernan and David Byrne

This double-act webinar was done by Mark Heffernan and David Byrne. You may have come across this duo at IATEFL if you attended. They also have a column in Modern English Teacher, who hosted this webinar. I haven’t encountered them before, but it was a really good webinar – if I were to attend an IATEFL in the future, I would totally look out for a session of theirs in the programme!

If you are they, or you attended the webinar, and see any mistakes in my notes-based summary, please comment and let me know!

The outline was as follows:

David particularly highlighted the idea of “Help your learners to find/make decisions”, saying that the role of teachers has changed over the years. We used to be arbiters of right and wrong, but now, we are facilitators of learning and discussion, our role isn’t to say what is right or wrong but to show possibilities and allow learners to make choices.

Writing

  • Has AI changed how we write?
  • Has AI changed how students write?

Yes.

Everyone (well, many people) uses it, to varying degrees of success, appropriateness and responsibility. If you don’t use it responsibly and effectively, it does wash out your personality/voice. In order to maintain your voice, you need to know what your voice is.

We have to train our learners on responsible, appropriate, effective use.

Questions we need to ask are: Who is the audience? What is the need (Why are you writing this?)? What role do you play in it? What role should/could AI play in this process?

E.g. a letter of complaint – if you will be all hedging/not cantankerous enough, you could use AI to write it and prompt it to add in some extra cantankerousness. If you are, you probably want your voice in there and will write it yourself. You have choices.

If we’re doing a test, AI is not appropriate unless it is built into the test. However, you could use it for brainstorming, ideation, feedback, suggested language chunks. It can be a learning tool. Most universities acknowledge and accept students using it in that way. What is generally prohibited is using it to produce text and submitting that. This is a change from two years ago and shows how things have evolved.

How do writers come across? How do you want to come across? It’s all about tone and voice.

The question becomes not did you get the grammar/vocabulary correct but is the text produced undeniably written by AI? If it is, it is not successful. If you have just pulled little language chunks from AI, then it could be.

You can teach a whole lesson on voice/tone but David/Mark suggest that is better to embed it throughout the course. Syllabuses tend to be spiral-shaped. Give students chances at multiple stages during the course to reflect and make choices. If we give them chances to do that, they have choices. It’s not a one and done lesson, appropriateness and AI can’t be a one off. It needs to be woven through. It needs to be scaffolded. The rise of AI has made it even more important than before to do this (teach about voice) but it was always important.

Speaking

When you speak, you portray a version of yourself, you make choices.

English learning and using depends on context: I need to be able to… so that I can… .

There is more than one correct way to structure an essay but we teach maybe the most foolproof way, the easiest way.

Hedging – it’s partly using modals, so it’s grammar but it’s also functional (you signal how sure or unsure, how strongly or otherwise you feel towards what you are saying).

David and Mark shared some possible activities for working with voice/persona by weaving it into existent activities:

If you don’t show interest in what someone is saying, so you just listen and don’t say anything/interject etc, the speaker may feel lack of interest and lose confidence. If you see this happen in a discussion between students of yours, facilitate discussion of these kind of moments – e.g. this happened (X didn’t say or do anything while you were talking), why is that, X? How did you feel about it Y?)

My take-away:

We have seminar discussion exam preparation and then the exams coming up, and I want to try taking this approach to evaluating the example discussion recording (e.g. how did x respond, or not, how do you think y felt?), and to feedback on students’ discussions, and link it back to the language we teach them in order to enable participation. Get them thinking about what kind of persona they want to portray in a seminar discussion exam (e.g. engaged, knowledgable etc) and how to achieve that, as well as get them thinking about how to participate effectively in a real seminar. I might get them to repeat a practise discussion while playing different personas, to give them a chance to experiment.

In terms of writing (we are about to embark on extended essay writing on Monday!), I want to include more discussion of voice and, again, showing them that they have choices over how to express themselves in their essays and how those choices affect the outcome.

I feel I’ve come away with a load of ideas for how to slightly tweak what I already do, and hopefully thereby increase the value of it to my students: I call that a win! 🙂 Thank you Mark and David!

Teacher Identity

This blog post was inspired by Sandy Millin’s write-up of an IATEFL 2025 panel on the subject of Teacher Identity

I think opportunities to discuss and reflect on teacher identity, such as the IATEFL 2025 panel written up by Sandy, are invaluable, as identity is constantly evolving and growing. In the first talk Sandy summarised, the speaker, Robyn Stewart, adapted Barkhulzen and Mandieta’s (2020) facets of language teacher professional identity to highlight the influence of the world on identity, external influence on it. It also shows the interplay between personal and professional identity and the elements that can be considered to be part of our professional identity:

Via Sandy Millin’s write-up of Robyn Stewart’s talk in the IATEFL 2025 teacher identity panel.

There are so many things that influence who we are in the classroom! One of the lessons Robyn Stewart drew from her dissertation research was “Don’t underestimate the role of context”. I’m inclined to agree:

On a personal level, I’m not that interested in generative AI, generally distrust it, disapprove of the resource consumption it represents and feel the amount of money, time, expertise and so on being ploughed into it everywhere could be better spent elsewhere (e.g. use of AI in medical contexts) rather than generating infinite quantities of text.

As a language learner, if I had the time, energy and spare brain, or was as driven as summer 2014 me, such that I could overcome the lack of all the afore-mentioned (and could override my concerns about unnecessary resource consumption!), I would perhaps explore the possibilities of communicating with it in Italian/French/German and using it to help me improve my production. I could get *well* in to a project like that. (And if I were teaching general English I could use the knowledge and skills I might develop in the process to help my students benefit from using the English version.)

However, my professional identity has the greatest influence on my interaction with AI: I have to embrace AI’s existence and figure out ways to work with students in a world which it is now very much a part of. In terms of context, I work specifically in higher education, preparing students to study at university by teaching them an Academic English skills course which they do alongside subject modules. Assessments are high stakes in terms of scores but they also need to ensure that students develop the skills necessary to succeed, including that of correctly treading the line between fair use of tools and academic misconduct regulations – a line that has been evolving with the evolution of AI. We used to mutter about Grammarly and translation tools, but ignore them other than prohibiting students from using them and putting a handful forward for misconduct each assessment cycle, and then generative AI came along and blew all that out of the water and onto a whole other level. We have been grappling with it ever since. However, it will only be come September of this year that I will engage with it fully as a teacher in the classroom beyond warning students off it (rather than only from the perspective of course coordination, course/materials development – as in, integrating teaching AI – related skills into our materials, currently in progress, rather than developing materials using AI – and misconduct evaluation).

The young Vietnamese participants in the study carried out by Hang Vu, the third speaker of the IATEFL 2025 panel on teacher identity, demonstrated a high level of insight and awareness into the issues they face in developing their professional identity as teachers in a world dominated by AI, and what kind of training they need in order to do that successfully. Sandy described Hang Vu’s idea of “emerging identities”, as summarised on the slide below:

Via Sandy Millin’s write up of Hang Vu’s talk in the IATEFL 2025 panel on Teacher Identity.

There’s a lot to think about there! I suppose I have mainly been teacher/coordinator as AI inspector in professional terms, but also teacher as learner as despite my personal misgivings: I have made an effort to attend (whether live or via recording) all the training available to us regarding AI. I have been teacher as AI user when I have used it to generate discussion questions (and then teacher as critical thinker when I have deleted half of them as unsuitable and edited/adapted others!). Teacher as AI instructor/facilitator, of course, as mentioned above, is still in the “coming soon to a classroom near you” stage. I suppose will have to be “teacher as AI supporter” within the “teacher as instructor/facilitator” side of things – regarding what we decide are acceptable uses of AI…but I predict it will be more along the lines of channeling inevitable use rather than encouraging use vs non-use! And I think alongside that, I will definitely be encouraging critical discussion in my classroom regarding the use of AI and surrounding issues. It will be interesting to see what the students think. It seems to me that just as much as the youngsters in the Vietnamese study, us old fossils who have been teaching a good while also need to regularly engage with our professional identities and figure out how we are going to move with the times professionally, regardless of (although obviously also interlinked/connected with/influenced by!) our personal feelings towards the various changes (which as Catherine Walters’ plenary discussed, have been many and varied over the last 50 years!)

Sandy’s post finished with some of the questions posed by the audience, one of which was “Should we proactively work with learners about how to do AI? Maybe we should ask learners for the whole AI conversation, not just the final result.” – It’s an interesting one. I definitely want critical discussion and to find out the students’ take on it, and as with other things potentially their feedback/ideas/thoughts can feed into future iterations of the course, but ultimately, in terms of assessment, what is and isn’t acceptable has to align with university and college policy on AI use. One thing I do hope is that I will be able to persuade students of the importance of developing their own voice, as I think if I can do that, then reasonable/acceptable use (with the appropriate guidance on how) will be a natural progression. For sure, all this thinking I am doing at the moment (I’m on annual leave – I have time to think!!) will be a useful form of preparation for the task ahead!

This blog post is plenty long enough already, yet I haven’t even scratched the surface of identity, personal and professional, and the interplays between identity and classroom. But, another time… 🙂

Generative AI and Voice

I’m a writer. I am writing right now! I have written journal articles, book chapters, (unpublished) fiction, (unpublished) poetry, materials, reflections (blog posts), combination summary/reflections of talks/workshops (blog posts) I attend, emails, feedback on students’ work, the occasional Facebook update, Whatsapp/messenger/Google chat messages, and so the list goes on. It is a form of expression, as is speaking, and drawing. These, including all the different kinds of writing I have done and do, are all forms of expression that AI is now capable of approximating. However, until fairly recently (when suddenly it was showing up everywhere!), I had not explicitly considered the relationship between AI generated production and a person’s ‘voice’. Examples of ‘voice’ vs AI can be seen in the two screenshots below:

Via an email from Pavillion ELT – abstract of a forthcoming webinar.
Via Sandy Millin’s summary of Ciaran Lynch’s MaW SIG PCE talk at IATEFL 2025.

Both of these screenshots set voice against AI-generated content. The first one (which looks like an interesting webinar – Wednesday 14th May between 1600 and 1700 London time in case you might like to attend!) seems to be about helping learners develop their own voice in another language and suggests that this aspect of language learning is of greater importance in a world full of AI output. The second is in the context of materials writing, and highlights an issue that arises in the use of AI in creating materials – “lacks teachers’ unique voice”. The speaker goes on to offer a framework for using AI to help with materials writing while avoiding the problems listed in the above screenshot. (See Sandy Millin’s write up for further information! The post actually collects all of her write-ups of the MaW SIG 2025 PCE talks in a single post – good value! 🙂 )

I teach academic skills including writing to primarily foundation and occasionally pre-masters students who want to go on and study at Sheffield University. In the last year, we’ve been overhauling our syllabus, partially in response to one of our assessments being retired and partially in response to the proliferation of generative AI. Our goal is to move from complete prohibition of AI to responsible use of it. And I suppose, one thing we hope to achieve from that is reach a point where students may or may not choose to use AI in certain elements of their assessment but actively avoid it in others. This, I think, has some overlap with Ciaran Lynch’s framework for writing materials:

Via Sandy Millin’s summary of Ciaran Lynch’s MaW SIG PCE talk at IATEFL 2025.

Maybe we need a similar framework/workflow for our students that succinctly captures when and how AI use might be helpful and when it is to be avoided. And I think voice is part of the key to that! But what exactly is voice? In terms of writing, according to Mhilli (2023),

“authorial voice is the identity of the author reflected in written discourse, where written discourse should be understood as an ever evolving and dynamic source of language features available to the writer to choose from to express their voice. To clarify further, authorial identity may encapsulate such extra-discoursal features as race, national origin, age, or gender. Authorial voice, in contrast, comprises, only those aspects of identity that can be traced in a piece of writing”.

[I recommend having a read of this article, if you are interested in the concept of voice! Especially regarding the tension between writers’ authentic L1 voice and the constraints of academic writing in terms of genre and linguistic features (which vary across fields).]

In terms of essay writing, and our students (who are only doing secondary research), if they are copying large chunks of text from generative AI, then they are not manipulating available language features to express meaning/their voice, they are merely doing the written equivalent of lip-synching. I think this is still the case if they use it for paraphrasing because paraphrasing is influenced by your stance towards what you are paraphrasing and how you are using the information. I suppose students could in theory prompt AI to take a particular stance in writing a paraphrase or explain how they plan to use the information but they would also need to be able to evaluate the output and assess whether it meets that brief sufficiently. In which case, would it save them much time or effort? Would the outcome be truer to the student’s own voice? I wonder. Of course, the assessment’s purpose and criteria would influence whether not that use was acceptable.

On the other hand, if students use AI to help them come up with keywords for searches and then look at titles and abstracts, and choose which sources to read in more depth, select ideas, engage with those ideas, evaluate them, synthesise them and organise it all into an essay, using language features available to them, then that incorporates use of AI but definitely doesn’t obscure their voice and the ownership of the essay is still very much with the student rather than with AI. They could even get AI to list relevant ideas for the essay title (with full awareness that any individual idea might be partly or fully a hallucination), thereby giving them a starting point of possible things to consider, and compare those with what they find in the literature. This (and the greyer area around paraphrasing explored above) suggests that a key element that underpins voice is that of criticality. Perhaps we could also describe it as active (and informed) use rather than passive use.

Another issue regarding voice in a world of AI generated output, which I have also come across recently lies in the use of AI detection tools:

From “AI, Academic Integrity and Authentic Assessment: An Ethical Path Forward for Education

If ESL and autistic voices are more likely to be flagged as AI generated content, then our AI detection tools do not allow space for these authentic voices. These findings point to a need to be very careful in the assumptions we make. I’m sure we’ve all looked at a piece of work and gone “this was definitely written by AI, it’s so obvious!” at some point. Hopefully our conclusions are based on our knowledge of our students, and their linguistic abilities, previous work produced under various conditions and so on. However, for submissions that are anonymised this is no longer possible. I think, rather than relying on detection tools, we need to work towards making our assessments and the criteria by which we assess robust enough to negate the need for such tools. Either way, the findings would also suggest that the webinar described in screenshot no. 1 may be very pertinent for teachers in our field. (I wonder if the speakers have come across instances of that line of research too?! I increasingly get the impression that schedule-willing, I may be attending that webinar!)

Finally, this excerpt from a Guardian article about AI and human intelligence I think provides perhaps the most important reason for helping students to develop their voice and not sidestep this through use of AI:

“‘Don’t ask what AI can do for us, ask what it is doing to us’: are ChatGPT and co harming human intelligence?” – Helen Thomson writing for The Guardian, Saturday 19th April 2025

We want those Eureka moments! We want the richness of what diversity of thought brings to the table. (It is baffling to see Diversity, Equality and Inclusion initiatives being dismantled willy nilly in the U.S. – everybody loses out from that. But then, so much of what goes on these days is baffling.) Maybe something small we can do is help our students realise that their voice, as every voice, is important and that diluting it and losing it through ineffective use of AI makes the world a poorer place. I haven’t even touched on AI and image production or AI and spoken production but this blog post is long enough already (maybe I should have got AI to summarise it for me! 😉 ) so I will leave that for another post!

Using Adobe Firefly for Image Generation

Have you used Adobe Firefly before? Me neither. But we have free access to it via the University and the TEL team has used it, and so did a session for us on it. It can be used to for images to put in lesson handouts and slides, but also online platforms like Wooclap and Quizlet.

You write a prompt in a box and it generates images.

This was a scenario given to us:

Prompt 1: an image of 4 students in a discussion. This was the result:

Issues: There are 3 students and teacher. They look quite young while we teach university age students. Three of them are blonde so it isn’t a good representation of our students. So this is an example of the bias that exists in AI in an automatic result with no detail prescribed in the prompt.

Prompt 2: an image of 4 university students from diverse background in a discussion. This was the result:

Problems: They are not in a classroom.

Adding “seated” (to be more typical of a classroom):

Not a perfect picture (looks a bit like an airport…) but better than the first picture! In terms of the purpose of generating the image, this would probably work. Prompt writing/editing for Adobe Firefly tends to take multiple iterations before you get something you might be happy to use.

We were given the following tips:

  • add more detail to get better results;
  • be aware of bias as you engineer prompts and evaluate the outcome;
  • be picky – it may take several iterations to get what you want. Sometimes a fairly simple prompt immediately yields a satisfactory outcome but usually it takes a bit more effort. Particularly to produce an outcome that is suitably representative for an international student population.

Adobe Firefly has a lot of stock images that it draws on which means the quality is better than similar counterparts.

Once you have generated an image you can also edit it to a certain extent. Which is good as the first images you get can have arm melds, funny shaped heads and so forth! It’s not very good with limbs. A central human image may be fine but anyone in the background or if you require groups/more people, then problems abound! Despite these issues, Firefly is better at it than Gemini.

So al very cool but actually stock images like Pixabay (and creative commons licensed like Flickr – in particular ELTpics – if the context is suitable), i.e. human generated, are much less resource-intensive to use. So, don’t get too carried away by the “it’s so cool” thing. I tend to use Google image search and the appropriate level of license filter, personally.

My general impression: I can’t currently see an Adobe Firefly – shaped hole in my life that needs filling. I wonder if in 5 years time I will look back on this post with an “oh you innocent child” type lens or not?! Time will tell! It was a good session though, after being shown the prompts and pitfalls, we went into a breakout group and had to come up with prompts for another scenario. Unfortunately in my group, none of us had access sorted out yet so we couldn’t test the prompts we wrote.