Evaluating Output in the AI Era

Welcome back to the second installment in our three-part series about how AI has (and hasn’t) changed LIB100. In our first blog post, we chatted about getting started with AI for research. In this blog post, we plan to discuss how AI has (and hasn’t) changed information evaluation, focusing on three main areas:

  • Hallucinations and Source Verification
  • Fact Checking AI Output: From Sources to Claims
  • Evaluation for Source Selection

A Mindset Shift

As discussed in the previous blog post, LIB 100 is a very popular credit-bearing course at ZSR Library and last academic year we taught nearly 320 students. When you assess the work of this many students year after year, shifts in their research behaviors are easily sensed. In addition to what we have perceived as changed research behaviors, students themselves are reporting changes that affirm our observations. In the years since generative AI tools have become widely adopted on our campus, we’ve been asking students several questions at the start of the semester including “What information skills do you think your generation is good at?” as well as “What information skills do you think your generation needs?” The majority of students report feeling that their generation is good at searching and finding information, mentioning Google, AI, and social media platforms as go-tos for “quick” information. The most common skills students believe their generation needs are related to verification, evaluation, and synthesis. This feedback and what we’ve noticed over the past few years has forced a major mindset shift from the classic “search-first” library instruction to meeting students where they are – in an “answers-first” information environment as UVa Library Dean Leo Lo has called it

Before generative AI was widely used by students, LIB 100 was created and taught from the assumption that students would not be able to find enough scholarly information to answer a research question for academic research. This approach was built on decades of helping students use digital platforms to access scholarly information at reference desks and in one-shots, and encouraged by the ACRL’s Framework for Information Literacy for Higher Education (“Searching as Strategic Exploration.”) Of course, an abundance of information existed for most of the student research topics that presented themselves during library instruction, but a certain knowledge of the structures organizing the material and how to effectively summon it was required to connect students to the information. While many students came to library instruction with a misguided idea of research (“I need articles to prove X…”), they still depended on the search phase for any subsequent research activities.

The current information environment our students inhabit has, at least from the average student’s perspective, upended the necessity of an advanced search process. Students may prompt a generative AI tool in exactly the way they used to come to the reference desk seeking articles to prove their assumptions, and in a matter of seconds what they asked for will be served up to them. LIB 100 is no longer taught from the assumption that students will have trouble finding information, but rather that they will find too much information, too easily and either trust outputs guilelessly or struggle to determine truth in the face of contradictory information on the same topic.

How we find information remains important, and we still devote considerable time to selecting appropriate search tools, choosing keywords, and applying Boolean operators and search filters. However, our new reality is that we must also prepare students to evaluate the answers they routinely encounter when Google or generative AI provide a synthesized response without the need to personally consult the underlying sources.

Hallucinations and Source Verification

So what does information evaluation look like in an “answers first” information environment? One of the first steps involves knowing that some answers may be based on fabricated or inaccurately cited sources – even if the claims are generally accurate. While most students know that AI can hallucinate facts, many still do not realize it can hallucinate scholarly sources. The frequency of these hallucinations varies by model and topic, with niche topics generally carrying a greater risk. That said, popular models are getting better, and we risk looking out of touch if we talk about hallucinations like it’s still 2023. When we started our “Source Verification” assignment 3~ years ago, roughly 30% of the citations that students generated in ChatGPT were completely fake. Another 30-40% were real but contained serious errors. Now, it’s much more likely that popular models will return real sources (<10% hallucination rate in our most recent semester) but with errors like incorrect DOI’s, volume and issue numbers, or author(s). Even though these models are getting better, we emphasize with this assignment that they are still far from perfect, and the news is rife with stories of where inaccurate AI-generated citations had significant professional and academic consequences. So, while spotting these errors can be tedious for both students and instructors, for now, at least, we are still teaching source verification as an important skill, typically by cross-checking the reference in Google Scholar and verifying the basic facts of the source. And though it’s tedious, the verification assignment comes up frequently in student evaluation comments about the most meaningful lessons learned in the course.

One of the ultimate goals of this exercise is to steer students towards better tools for information discovery. ChatGPT and Claude have their uses, but currently, better tools exist for finding high-quality academic information. In addition to the library home page (Primo), library databases, and Google Scholar, we’ve begun teaching students how to use research-specific AI tools like Scite, Consensus, and Google Scholar Labs. While these tools have their drawbacks (which we’ll discuss more in part three), unlike popular generative AI models, they search verified academic indexes, reducing if not completely eliminating the hallucination problem. Learning about AI research tools is hands down the part of the course that receives the most positive feedback from students.

Fact Checking AI Output: From Sources to Claims

While we just spent two paragraphs discussing the importance of verifying the existence of sources, in an answers-first environment, one of the biggest shifts in thinking involves moving from primarily source-focused evaluation methods to claim-focused evaluation methods. In the pre-AI era, evaluating information focused largely on the idea that authority is traced back to identifiable sources. We could generally rely on criteria like the expertise of the author and reputability of the publication to give us clues about the reliability of the information presented. But what happens when, as it often is with popular generative AI models, the author or organization responsible for the information is no longer identifiable? Thankfully, Mike Caulfield’s work with the SIFT Method gave us a good jumping off point discussing how to evaluate claims rather using lateral reading techniques – which to be fair, often rely on cross-checking those claims in reputable, identifiable sources. In addition to looking for other reputable coverage on the claim, we also emphasize tracing claims and quotes to their original sources.

SIFT has even been automated through Caulfield’s SIFT Toolbox for AI, which is a prompt Caulfield developed for general use GAI tools like Claude or ChatGPT, and which can be copied and pasted into your tool du jour as a workflow to test viral claims, images, or headlines. When we piloted this (with small groups of LIB 110: Research After Wake students, initially) it was an excellent shortcut for students wanting to test some of the claims from social media content. However, as we started using this prompt with a greater volume of students, we noticed that the tools were generating inconsistencies when provided with the same ‘claim’ to verify both across tools as well as within the same tool for different users. Of course, this can be an excellent teaching exercise in itself by asking students to carefully compare the outputs and different metrics, source materials, and logic used to determine ‘correctness,’ but as with using GAI tools for the previous research tasks it left us asking whether or not the ‘shortcut’ was really saving us any time at all.

That said, AI has made other aspects of lateral reading more complicated (e.g., students may mistake reading Google AI summaries for lateral reading, when it often summarizes the very source they are trying to fact-check), and there are times when we still do need to evaluate sources rather than claims.

Evaluation for Source Selection

Finally, though students have easy access to information, are they finding the best information? And do they have the skills to know what high quality information is? While selecting the best information available has always been a goal in LIB100, AI has made it more complicated. For example, we have never seen students find so many arXiv pre-prints than we have in the last two years, often without recognizing that these papers may not yet made it all the way through peer review. Five years ago we might have said talking about pre-prints was “more information than students really needed,” but times change. Additionally, unlike library databases, which present individual studies for users to examine, AI research tools often summarize findings in a uniformly confident voice that can obscure meaningful differences in study quality across different sources. This is especially pronounced in areas where high quality research is thin. This can lead students to assume that all the studies included in an AI-generated summary carry equal authority, even when some rely on weak study designs.

Teaching students to think about study quality is a fairly advanced and nuanced skillset for LIB100, but we introduce concepts like methodological rigor, scholarly influence, and alignment with scholarly consensus so that students can begin to interrogate study quality at an introductory level. Surface-level measures, such as journal rankings and citation counts, can be blunt instruments with problematic implications, so we are careful to highlight how reliance on these measures can perpetuate epistemic injustice by privileging established scholars, institutions, journals, and ways of knowing while marginalizing others. That said, these measures can still provide useful context when treated as starting points.

Finally, we invite students to consider what’s available to be searched in AI research tools – the answer is different in almost every tool, largely due to licensing agreements. While we’ve always emphasized that no search tool can find everything, and that good search strategies involve searching in multiple places, we are making a point to emphasize the gaps in scholarly coverage that exist in AI research tools. Compared with the greater effort required to search a library database effectively, the ease with which AI returns results can create the false impression that those results represent the full body of available research. Recognizing these gaps helps students treat AI-generated results as one entry point into the scholarly conversation rather than a comprehensive account of it.

Conclusion

Ultimately, how we evaluate information is contextual. We may already intuitively realize this – for example, how we evaluate social media content differs in some ways from how we evaluate scholarly articles. Similarly, how we evaluate claims may differ in some way from how we evaluate identifiable sources, and how we might evaluate output from a general AI model like ChatGPT or Claude has differences from how we might evaluate output from an AI tool that’s searching from a verified academic index. Holding all those threads together can be difficult, and the material realities of most library instruction mean we often boil down information evaluation into 4-5 discrete steps with cute acronyms like CRAAP and SIFT. One of the many benefits of a credit course is that we have time to emphasize the contextual nature of evaluation and help students begin choosing strategies suited to the particular information they encounter.

That was plenty for now. In our final blog post we will discuss even more about evaluation – from evaluating AI summaries and synthesis to re-thinking AI research projects in the AI era.