Sally Christopher
DVM

Artificial intelligence (AI) platforms entered the field of veterinary radiology about 10 years ago. And as we all know, the abilities of AI have expanded enormously over the past year. But how accurate are these new algorithms? Proprietary datasets are used to train these platforms, and little to no information is disclosed regarding how and when the algorithms are updated. Independent external validation is needed to evaluate the accuracy of commercial veterinary radiology AI platforms.
Editorās Note: This is an excerpt from Research Wrapped, a free monthly newsletter that collects the latest scientific research relevant to small animal veterinarians and pulls out practical takeaways. To be the first to receive this newsletter each month, subscribe here.
In a recent pilot study, published online ahead of print inĀ JAVMA, researchers aimed to do just that by assessing 6 AI platforms (Picoxia, Radimal, RapidRead, SignalRAY, Vetology AI, and X Caliber Vet AI) offering veterinary radiography interpretation services. (The study was not designed to specifically compare the 6 AI platforms.) The data includes abdominal radiographs of 53 dogs with definitive diagnoses. Researchers organized the data by confirmed diagnoses (e.g., small intestinal obstruction, pneumoperitoneum, gastric dilatation volvulus). Ultimately, the researchers concluded that all 6 AI platforms are not qualified for clinical diagnostic purposes as indicated by frequent missed diagnoses.
I spoke with 2 of the authors,Ā Doris Ma, BVSc, MVetClinStud, andĀ Stephen Joslyn, BVSc, BVMS, DECVDI, about their thoughts on this study and the applications of AI in veterinary medicine.
Q: Considering there are only a few published studies on the performance of AI, how did you select the studyās design and testing approach?
We deliberately chose an external validation design using real world cases sourced from general practice rather than curated or idealized datasets. Each case had a definitive diagnosis established through surgery, advanced imaging, histopathology, or clinical outcomeāallowing for meaningful ground truth comparison.
Our goal was not to reflect how these platforms are used in practice. We prioritized metrics such as Matthews correlation coefficient (MCC) and balanced accuracy, which better account for class imbalance and provide a more reliable assessment than accuracy alone.
Q: The results of the study support the general concerns of overreliance and inconsistent performance of AI and suggest that the current algorithms are not equipped to independently assess abdominal radiographs in clinical practice. Would you consider that to be the primary takeaway of this study?
Yes, that is a fair summary. The primary takeaway is that current commercial AI platforms demonstrate variable and, at times, limited performance when applied to real world cases.
Although these tools may have potential as adjunctive aids, our findings suggest that they are not yet reliable enough for clinical use. This is particularly important given the risk for overreliance, especially in settings where users may not be fully aware of the limitations of the algorithms.
Q: You list several limitations of the study, including potential selection for complex cases within the dataset, observer bias, small sample size, and imbalanced dataset (i.e., very few normal cases). How do you recommend future studies avoid imbalanced datasets to better analyze AI algorithms?
Future studies should aim for more balanced datasets by prospectively recruiting cases to ensure adequate representation of both normal and abnormal studies.
Editorās Note:Ā This article is an excerpt from the Research Wrapped monthlyĀ newsletter.Ā Subscribe here for free.
Q: What do you foresee to be the greatest areas in which AI may contribute to the advancement of veterinary medicine research?
AI has significant potential in areas beyond stand-alone image interpretation. In particular, integrating images with clinical history, signalment, and longitudinal patient data may enable more meaningful clinical decision support. But we currently have insidious issues with animal health data. Itās full of errors, assigned to wrong patients, missing data (multiple clinics), and no one uses a unique identifier!
Q: Is there anything you would like our readers to know that has not been mentioned?
A key consideration is the need for transparency and reproducibility in AI evaluation. Many commercial platforms do not disclose training data, model updates, or versioning, which makes independent assessment challenging. We would encourage the development of standardized external validation frameworks and greater collaboration between clinicians, researchers, and industry. This will be essential to ensure that AI tools are safe, effective, and appropriately integrated into clinical practice. On that note, Vetology AI has recently released their performance metric. This is a great initiative, and I hope other service providers follow this transparency practice.
Lastly, clinicians should be cautious when interpreting marketing claims around āstamps,ā āwarranties,ā or āassurances.ā While these published statements may suggest that AI providers will support practitioners in the event of an incorrect diagnosis, in practice, the duty of care remains with the treating veterinarian and such assurances do not equate to meaningful clinical, financial, or legal protection. It is therefore essential that these tools are used with appropriate clinical oversight and not relied upon as a substitute for professional judgment. Framing AI as simply āan extra pair of eyesā may understate the responsibility associated with acting on its outputs; clinicians must understand the limitations and performance characteristics of these systems. For example, if an AI algorithm suggests a diagnosis such as small intestinal obstruction and this directly influences a decision to proceed with exploratory surgery, the responsibility for that decision remains with the clinician. In the event of an incorrect outcome, this may reasonably lead to client dissatisfaction, financial disputes, or complaintsāirrespective of the AI input.
Read the Full Study
Pilot study: external validation of commercial veterinary radiology artificial intelligence services shows deficiencies in interpretation of general practice-sourced canine abdominal radiographs.
Ma D, Faulkner JE, Stander N, Raisis A, Joslyn S. JAVMA | Published online March 20, 2026 | doi:10.2460/javma.25.10.0691
