← Latest papers
💻 computer science

Coligo: A Retrieval-Augmented Generation Assistant for WhatsApp-Based TNEA Engineering Admission Counselling

Coligo is a Retrieval-Augmented Generation assistant deployed on WhatsApp that alleviates the administrative burden of Tamil Nadu Engineering Admissions (TNEA) counselling by leveraging Google Gemini and vector search over college documents to provide accurate, context-aware, and category-specific admission guidance while explicitly distinguishing between its functional prototype and its intended full-scale architecture.

Original authors: Nithishkumar A, Nitin A V, Prasanna S

Published 2026-09-22
📖 1 min read☕ Coffee break read

Original authors: Nithishkumar A, Nitin A V, Prasanna S

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: Coligo – A Retrieval-Augmented Generation Assistant for WhatsApp-Based TNEA Engineering Admission Counselling

Problem Statement
Engineering admission counselling in Tamil Nadu (TNEA) creates a bottleneck for college front offices, which are inundated with repetitive queries regarding admission procedures, branch-wise fees, and community-specific cutoff ranks (OC, BC, BCM, MBC, SC, SCA, ST). Current solutions rely on manual staff intervention or generic chatbots that lack semantic understanding and access to verified institutional data. A critical gap exists in providing answers that are:

  1. Grounded: Based strictly on official college documents rather than model hallucinations.
  2. Community-Aware: Distinguishing between different reservation categories and cutoff types (marks vs. ranks).
  3. Accessible: Available on the platform students already use (WhatsApp) without requiring new app installations.
  4. Contextual: Capable of handling follow-up questions (e.g., "What about ECE?") without requiring the user to restate the full context.

Methodology and System Architecture
Coligo is a Retrieval-Augmented Generation (RAG) system designed as a two-component Docker Compose stack (FastAPI service and PostgreSQL database). The system architecture is as follows:

  • Ingestion Pipeline: Official college PDFs (merging admission procedures, academics, and cutoffs) are processed via pypdf. Text is extracted and split into overlapping chunks (512 words with a 50-word overlap) to preserve context across boundaries. Chunks are hashed (SHA-256) to prevent duplicate embedding during re-ingestion.
  • Vector Storage: Chunks are embedded using Google's text-embedding-004 model (768 dimensions) and stored in a PostgreSQL database extended with pgvector. This allows for semantic similarity search using cosine distance.
  • Retrieval and Generation:
    • Query Rewriting: To handle follow-up questions, the system uses a session-based memory (keyed by WhatsApp phone number or session ID). If a query is context-dependent (e.g., "What about ECE?"), a separate LLM call rewrites the query into a standalone form for retrieval purposes, while the original query is retained for the final answer generation.
    • RAG Pipeline: The system retrieves the top 5 (RAG_TOP_K=5) relevant chunks. These are combined with a strict system prompt that enforces domain rules: distinguishing cutoff marks from ranks, responding only to specific reservation categories, and refusing to answer if the context is insufficient.
    • Generation: Google Gemini (LLM) generates the final response grounded only in the retrieved context.
  • WhatsApp Integration: The system connects via the WhatsApp Cloud API. It implements webhook signature verification (X-Hub-Signature-256) and routes messages based on sender identity: student queries go to the RAG pipeline, while admin queries are routed for command processing (though command execution is currently limited).
  • Deployment: The stack is containerized using Docker Compose, with specific configurations to handle DNS resolution issues in containerized environments (pinning external resolvers).

Key Contributions
The paper explicitly distinguishes between the working prototype and the originally planned architecture, highlighting the following implemented contributions:

  1. Containerized TNEA Pipeline: A working RAG system deployed on WhatsApp that relies on official PDF ingestion rather than model memory.
  2. Domain-Specific System Prompt: A prompt engineered to handle TNEA-specific logic, including the cutoff calculation formula (Math/2 + Physics/4 + Chemistry/4) and the necessity of community-wise responses.
  3. Resilient Conversation Memory: A session-scoped memory design that degrades gracefully to stateless answers if database history access fails, rather than crashing.
  4. Reproducible Deployment: A documented Docker Compose setup for PostgreSQL with pgvector and FastAPI, including fixes for container DNS resolution.
  5. Transparent Reporting: An explicit, code-verified account of which architectural components (e.g., Redis caching, Celery workers, multi-provider LLM routing, admin command execution) are not yet implemented, preventing the misrepresentation of the prototype as a full production system.

Results and Verification
The authors verified the system by rebuilding the stack from a clean state and running ingestion code against a 9-page, 5,048-word PDF from Sri Krishna College of Engineering and Technology (SKCET).

  • Ingestion: The pipeline successfully processed the PDF into 11 overlapping chunks, with the first chunk at 512 words and the last at 428 words, confirming the chunking logic works as intended.
  • Deployment: The Docker Compose stack started successfully, with health checks (/health and /health/detailed) returning HTTP 200 and confirming database connectivity.
  • Limitations in Verification: Due to the absence of a configured GEMINI_API_KEY in the verification environment, live latency, retrieval accuracy, and answer-grounding metrics were not measured numerically. The "retrieval and generation" phase was validated via code review rather than live execution.
  • Admin Routing: While the system correctly detects and logs admin commands (e.g., /ingest, /stats), the actual execution of these commands is not yet implemented.

Significance and Claims
The paper positions Coligo not as a finished commercial product, but as a verified prototype that bridges the gap between generic LLM chatbots and rigid keyword-based systems. Its primary significance lies in:

  1. Honesty in Engineering: The authors explicitly document the "gap" between the designed architecture (which included Redis, Celery, and multi-LLM routing) and the current implementation. They argue that distinguishing the working prototype from the target architecture is a contribution in itself, ensuring stakeholders do not mistake the current state for the final design.
  2. Practical Accessibility: By utilizing WhatsApp, the system removes the barrier of app installation for students and parents, meeting them on a familiar platform.
  3. Grounded Reliability: The system prioritizes factual accuracy over fluency, explicitly programmed to refuse answering if the source document does not contain the specific community-based data, thereby reducing the risk of hallucinated admission cutoffs.

The paper concludes that while the core ingestion and retrieval pipeline is functional and verified, future work is required to implement admin command execution, confidence scoring, human-in-the-loop escalation, and multi-LLM routing to realize the full scope of the proposed architecture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →