← Latest papers
🤖 machine learning

VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening

VetClaw is an edge-cloud multimodal agentic system that integrates an edge-based interaction manager with a cloud-hosted LangGraph workflow to enhance veterinary disease screening by combining visual and textual inputs for improved zero-shot classification while ensuring safety, reliability, and structured diagnostic support.

Original authors: Syed Mhamudul Hasan, Anas AlSobeh, Hussein Zangoti, Abdur R. Shahid

Published 2026-07-29
📖 5 min read🧠 Deep dive

Original authors: Syed Mhamudul Hasan, Anas AlSobeh, Hussein Zangoti, Abdur R. Shahid

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where your pet could talk. If your dog had a sore paw, it wouldn't just limp; it would say, "Hey, my paw hurts, and it looks red." In reality, animals can't speak our language, and they can't fill out medical forms. This silence makes it hard for veterinarians to catch diseases early, especially when a sick animal can't tell you what's wrong. To bridge this gap, scientists are building "smart helpers" using Artificial Intelligence (AI). Think of AI as a super-observant detective that can look at a picture of a sick animal and read a description of its symptoms to guess what's wrong. But just having a detective isn't enough; you need a whole team. You need someone to take the photo, someone to organize the clues, someone to check if the detective is making things up, and someone to call for help if the case is too tricky. This is where the concepts of "edge computing" (doing the work right where the animal is, like on a small camera) and "agentic systems" (AI that acts like a team of workers rather than just a single calculator) come in. The big question is: Can we build a digital team that works together to spot animal sickness before it gets serious, even if the animal can't say a word?

Enter VetClaw, a new digital system designed to be that very team. Imagine VetClaw as a high-tech "pet health detective squad" deployed on a tiny computer called a Raspberry Pi, which acts like a smart camera on a farm or in a barn. This system doesn't just snap a photo and guess; it works like a well-oiled machine with two main parts. First, there's the "field agent" (called OpenClaw) that lives on the camera. It's the one that wakes up, takes a picture of a cow or a dog, and asks the human nearby, "Does this animal seem sick? What do you see?" It collects the photo and any notes the human writes down. Then, it hands these clues to the "mission control" (called LangGraph). Mission control is the brain that organizes the workflow. It checks if the photo is clear, sends the clues to a powerful computer in the cloud, and waits for the answer.

The cloud computer uses a special type of AI called a Vision-Language Model (VLM). Think of this model as a super-smart veterinary student who has read every medical book but has never seen a real animal before. It looks at the photo and the notes to guess the disease. But here's the twist: VetClaw doesn't just trust this student blindly. The "mission control" has a strict set of safety rules. If the student seems unsure, or if the photo is blurry, or if the notes say something urgent, the system doesn't just shout out a diagnosis. Instead, it flags the case, logs the details, and sends a "safety alert" to a human vet to take a closer look. It's like a junior detective who knows when to say, "I think I know what's wrong, but let's get the Chief to double-check this."

The researchers tested this system using two sets of pictures: one with 1,736 images of various pet diseases and another with 443 pictures of dogs with skin issues. They tried three different ways to feed information to the AI: just the picture, just the text description, and a mix of both. The results were fascinating. When the AI only looked at the pictures (like trying to guess a disease from a blurry snapshot), it wasn't very good at it. For the dog skin disease test, it only got about 33% of the answers right. However, when the AI was allowed to read the symptoms and look at the picture, its performance jumped significantly. In the dog skin test, the accuracy rose to over 72%, and in the broader pet disease test, the mix of text and images helped the system make much smarter guesses. Interestingly, in some cases, just reading the text description of the symptoms was even more powerful than looking at the picture alone, suggesting that knowing what to look for is often more important than just seeing it.

The paper suggests that this "team approach" is the future of animal health monitoring. It shows that combining a camera, a human's observations, and a smart AI workflow can catch diseases earlier than a simple camera ever could. However, the authors are careful to say this is a "work-in-progress." They admit that their tests used relatively small groups of pictures and that the AI was guessing without having been specifically trained on thousands of sick animals (a method called "zero-shot"). They also note that the current camera can only see from one angle, which might miss things if an animal is moving. But the core idea is solid: by turning a static camera into a coordinated, safety-aware team, VetClaw suggests we can build a system that doesn't just predict sickness, but manages the whole process of checking, verifying, and calling for help when needed. It's a step toward a world where our furry friends get the care they need, even when they can't speak up for themselves.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →