Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

Wednesday, August 26, 2026

OpenEvidence quietly makes accounts available to medical librarians

OpenEvidence is a large language model (LLM) grounded in the medical literature, with its website self-identifying as an "AI copilot for doctors that helps them make high-stakes decisions at the point of care." The free tool has become increasingly popular within the health sciences, with its website boasting that "This year, more than 100 million Americans will be treated by a clinician using OpenEvidence."

Up until now, OpenEvidence accounts have been restricted to health care professionals with a National Provider Identifier (NPI) or students that provide proof of enrollment to a health sciences college or university, which has made evaluation of the tool particularly challenging for librarians (but even that hasn't stopped us: Exhibits A and B. Librarians are a tenacious bunch!😉).

However, at long last (and likely in consequence of tenacious librarians spamming their inboxes),  OpenEvidence has quietly opened account access to medical librarians. Now, librarians have the option to register as a "medical librarian" so long as they upload proof of their current position title and employment (I was able to create my account by uploading a screenshot of my public staff page, which includes my photo and my position as a health sciences librarian). Not only does this give you unlimited queries (without an account you were only allotted 2 or 3 queries per week), but it gives you access to additional tools in OpenEvidence, such as their calculators, clinical trials locators, etc.

For those who have seen one of our Wild West presentations or our MLA poster session, or read our article on OpenEvidence, you've likely gotten the impression that my colleagues and I are more than a little skeptical of OpenEvidence; and we've actively encouraged others to closely evaluate its outputs.

However, giving credit where credit is due, OpenEvidence opening access to medical librarians is a step in the right direction, and will allow us to get an inside look into the tool. 

Now that we've got the keys to the castle, we can better search for cracks in the foundation. 🔍

Castle

Happy sleuthing, everyone! ☕

Friday, July 10, 2026

2026 NLM/MLA Joseph Leiter Lecture: How AI is Reshaping Biomedical Discovery


 

Monday, June 15, 2026

Article of Interest: Limitations of Open Evidence (by PAIJE!)

Congratulations to Paije Wilson and her four colleagues from UW–Madison on their recent publication in the Journal of the American College of Clinical Pharmacy!

Their article tackles the critical topic of AI accuracy in healthcare, providing a thorough review of current literature on OpenEvidence. By running specific pharmacotherapeutic queries, the team effectively demonstrated inaccuracies in how the tool generates responses and summarizes its sources. Crucially, the authors delve into source summarization inaccuracies—a nuance in AI performance that has largely been overlooked in current research.


Link to study: https://accpjournals.onlinelibrary.wiley.com/doi/10.1002/jac5.70237

Wednesday, April 22, 2026

Misinformation in artificial intelligence tools: The game is afoot!

Magnifying glass

Image by Markus Winkler from Pixabay

I came across an interesting (and rather alarming!) read from Nature.

In the study, researchers posted two papers to a preprint server discussing a fake disease called Bixonimania, with the purpose of seeing whether existing large language models (LLMs) would reference the papers in its health advice. The researchers included multiple "tips" in the papers' full text identifying them as fake (my favorite was an acknowledgment to someone from the Starfleet Academy!). 

Despite these obvious tips, not only were the papers cited in LLMs' generated summaries, but were cited by a few peer reviewed publications as though they were legitimate sources! The researchers deduced this latter result may be attributed to authors' relying on AI generated references for their research without reading the full text.

This study illustrates not only the dangers of relying on LLM-generated summaries for advice (especially when that advice is medical!), but also relying on these summaries for generating citations for one's research. 

Even AI literature summarizers that are supposedly dedicated to academic and medical research are subject to these pitfalls. Myself and my colleagues at the Ebling Library have compiled several examples of such AI tools citing lower quality studies, and, in many cases, wholly misrepresenting the contents of the articles they cite.

As those who have read about my previous clown shenanigans are all too aware (here are my first and second blog posts on the topic, if you would like some humorous reads!), even AI tools designed to "read" full text PDFs don't always pick up on obvious red flags, and can misrepresent the contents of an article. 

As librarians, catching AI in these errors can feel a bit like detective work; however, what with all the hype relating to AI in research, alerting researchers to the current limitations of these tools is essential. As Sir Arthur Conan Doyle's Sherlock Holmes would say, "The game is afoot!"

Friday, December 19, 2025

Blog post with tips for spotting hallucinations in AI generated content

Spotting scope

Image by Afif Ramdhasuma from Pixabay

A new post published December 16, 2025 on the blog, Card Catalog, discusses practical tips for identifying hallucinations in AI generated citations. The post, titled, How to spot AI hallucinations like a reference librarian by Hana Lee Goldin, provides a quick, plain-language overview of why AI has the tendency to hallucinate references, and some tell-tale signs of hallucinated content. 

Something I particularly appreciate is that, in addition to providing tips for determining if a citation exists, the post also provides tips for verifying whether the AI is accurately summarizing the sources it's citing, being a vital check that often gets overlooked in AI generated content.

Happy reading, and hope everyone has a great weekend! ⛄

Thursday, November 6, 2025

Open Evidence-Krafty Librarian

 

Did you catch the Krafty Librarian's post, "OpenEvidence: Smart Medicine or Smart Marketing?"  After meeting our new Internal Medicine Program Director, who had questions about it, I’ve been dabbling with OpenMedicine myself. While Michelle reports accurately that you need an NPI to create an account; I just used my hospital's NPI, and I was in without issue. (If you want a way around that)

The big question, of course, is its utility. I recently leveraged it not as a substitute for a comprehensive search, but to add to one. Specifically, I used it to reinforce my search results on best practices literature, giving me a quick double-check on established evidence to ensure I had a complete picture.

The original blog post raises vital questions about balancing slick presentation with true evidence integrity, and it challenges us to place resources like OpenEvidence in the correct context for our users. Is it a time-saver? Yes. Is it a perfect primary source? Probably not. We need to be the critical thinkers guiding our clinicians and researchers through the noise. 

https://kraftylibrarian.com/openevidence-smart-medicine-or-smart-marketing/ 

Monday, July 14, 2025

Clowning around with AI: Experimenting with article PDF summarizer tools

Colorful assortment of balloons

There has been an explosion of artificial intelligence (AI) tools over the past few years. A category of AI tool that has been getting some traction is tools that summarize individual articles. A few such tools include Elicit, SciSpace, Perplexity, and EndNote 2025's new Key Takeaway tool (however, there are many, many more out there!). Among other things, these generative AI tools provide brief, easily digestible summaries of an article.

While these article summarizers have been lauded for their efficiency, there have been some concerns relating to their accuracy. Additionally, with the black box nature of AI tools, it can be difficult to determine just how much of an article AI tools are "looking at" when generating high level summaries.

Enter the Clown Shenanigans 

As someone who recently got access to EndNote 2025's Key Takeaway tool, I decided to play around (or, more aptly, clown around) with the tool. Using the text of an article I had published with JMLA, I systematically replaced different sections of the article with nonsense text to see if the Key Takeaway tool would pick up on the shenanigans. The "nonsense text" consisted of snippets of a fictitious study on identifying malicious clowns hiding within the general public, which I generated using Microsoft Copilot.

Findings 

In terms of replacing individual parts of an article, one section at a time, with clown nonsense, I found each and every replaced section (i.e., title, abstract, introduction, methods, results, discussion, conclusion, and references), by itself, managed to fly under the radar in the Key Takeaway tool (i.e., no clown shenanigans detected).

I also tested a few section combinations. My most interesting finding was that I was able to fully replace the methods, results, and references sections of the article (resulting in 47% of the article being comprised of text about clowns) without EndNote 2025's Key Takeaway tool mentioning anything about clowns in its generated summary!

Screenshot of EndNote's Key Takeaway tool. The methods section of the PDF has been replaced by clown nonsense, and the generated summary doesn't mention any clowns in it

I tested this same PDF (i.e., with nonsense methods, results, and references) out in SciSpace, Perplexity, and Elicit, and the clown shenanigans remained undetected in their generated summaries, as well (note that I only tested the high level summaries, and not the summaries these tools generated for each individual section of the article).

Takeaways

This fun little experiment only further illustrates the need to take caution with these AI summarizer tools, especially those that generate high level summaries. Though these tools can be handy, they can sometimes miss much needed context (or, in this case, clown shenanigans!) that may be buried in the full text of an article. 
 
While I would hope authors wouldn't replace entire sections of their manuscripts with nonsense, the fact that I was able to wholly replace vital sections of the manuscript, such as the methods section, without the text making its way into the high level summaries demonstrates how researchers relying on such summaries may miss necessary context or, perhaps most concerning, severe methodological flaws, if they don't take the time to read the studies in their entirety. To be fair, though, this is true of any high-level summaries, not just AI outputted ones.
 
Generative AI tools are ever evolving, and issues such as these may (hopefully!) be soon resolved. In the meantime, I encourage others to clown around with these tools to explore their strengths and limitations. For those wanting to conduct experiments of their own (or wanting a good laugh), here is a link to the different sections of text generated by Copilot (note that I didn't author any of the text, and the text is wholly the output of Copilot). See which sections of an article you can replace! Detection of clown shenanigans may vary.
 
Thanks for reading, and I hope everyone has a great week!