AI Image Errors in Science Are Getting Harder to Spot — So the Review Process Has to Change Too
The problem is no longer just “fake images”
A lot of the discussion around AI-generated scientific figures assumes the main risk is deliberate fraud. But many of the current failures are quieter than that: mislabeled panels, subtly altered microscopy textures, impossible shadows in conceptual diagrams, or AI “cleanup” tools changing raw experimental data without the author realizing it. Nature’s recent guidance is notable because it focuses less on punishment and more on process discipline. The recommendations are simple: know the publisher’s rules, never manipulate original data, and take responsibility for every figure that appears under your name.[1]
That sounds obvious until you look at how fragmented current publishing policies are. Different journals draw different lines around AI-assisted image editing, enhancement, and generation.[3] Some allow limited cosmetic edits. Others prohibit AI-modified research images entirely.[5] Researchers moving quickly between preprints, conferences, and journals can easily assume a workflow is acceptable when it is not.
The interesting shift here is cultural: image integrity is becoming less about “did a human make this?” and more about “can the provenance of this figure be reconstructed?”
Scientific images now need something like version control
Software teams learned this lesson years ago. You do not just ship a build; you keep a record of what changed, when, and why.
Science images increasingly need the same treatment.
Western blot guidance and journal integrity policies already emphasize retaining original files and documenting adjustments.[5] Nature’s advice similarly centers on preserving raw data and being accountable for edits.[1] That becomes much harder when generative tools are embedded directly into editing software. A researcher may use an AI feature for denoising or object removal without realizing the model synthesized entirely new visual information.
The technical issue is that generative systems are probabilistic. They are designed to create plausible pixels, not historically accurate ones.
That distinction matters in science more than almost anywhere else.
Detection tools are not a reliable safety net
There is a growing assumption that AI detectors will eventually solve this. The evidence so far suggests otherwise.
Research on scientific integrity tools notes that AI detection systems remain limited and are engaged in a constant feedback loop with tools designed to evade them.[6] Similar frustrations already appear in education, where both instructors and students describe unreliable detection systems and behavioral distortions caused by them.[2][4]
That has an important implication for scientific publishing: forensic detection alone is probably the wrong layer to depend on.
Instead, journals and labs may need workflows that assume AI assistance is common and focus on traceability: - retaining raw acquisition files - documenting transformations - separating illustrative art from evidentiary images - requiring disclosure of generative tools - making figure review part of peer review rather than a post-publication integrity check
That approach is less dramatic than “AI detector catches fake paper,” but it is probably more realistic.
The deeper issue is trust, not aesthetics
The most convincing AI-generated science image errors are often visually impressive. That is part of the danger. Humans are very good at mistaking coherence for accuracy.
Scientific publishing has historically relied on a mix of institutional trust, peer scrutiny, and reproducibility. Generative AI stresses all three because it lowers the cost of producing polished-looking material at scale.[3]
What makes this moment interesting is that the solution may not come from better image generation or better detection alone. It may come from adopting more transparent production practices borrowed from software engineering, open science, and archival workflows.
In other words, the future of trustworthy scientific imagery may depend less on whether AI was used — and more on whether every meaningful change can still be audited afterward.[1][5][6]
Sources
- 1] Three tips to avoid AI image mistakes in science - Nature — [https://www.nature.com/articles/d41586-026-02233-w
- 2] Students are deliberately writing worse to avoid AI detection flags ... — [https://www.reddit.com/r/Professors/comments/1rjl5u0/students_are_deliberately_writing_worse_to_avoid/
- 3] What do academic publishers say about using gen AI? — [https://janetsalmons.substack.com/p/what-do-academic-publishers-say-about
- 4] I used the hidden white text method to detect AI. And had to report ... — [https://www.reddit.com/r/Professors/comments/1p0m66o/i_used_the_hidden_white_text_method_to_detect_ai/
- 5] From Gel to Journal: A Practical Guide to Western Blot Submission — [https://www.bioradiations.com/from-gel-to-journal-a-practical-guide-to-western-blot-submission/
- 6] AI for scientific integrity: detecting ethical breaches, errors ... - PMC — [https://pmc.ncbi.nlm.nih.gov/articles/PMC12436494/

