Inspigo AI: Real-Time RAG for Personalized Learning
A streaming chatbot experience powered by OpenAI/Bedrock and a RAG pipeline, grounded in the platform’s own learning content.
Problem
Personalized learning benefits from conversational, context-aware assistance rather than static content lookup — but a foundation model answering purely from its general training data isn't grounded in what the platform actually teaches.
Architecture
Source documents are uploaded through the Inspigo CMS, stored in S3, and ingested into Pinecone as the vector store — Pinecone handles chunking and metadata, with embeddings generated via OpenAI's embedding models. At generation time, retrieved chunks are injected into the prompt before the model responds. OpenAI is the primary model, with AWS Bedrock as a fallback.
Execution
Owned the frontend/streaming UX end to end, plus the CMS interface used to upload and manage source documents. Streaming shipped first over WebSocket (still powering some legacy roleplay features), with newer flows moved to SSE. The main hard problem was scaling the streaming layer under concurrent load — getting real-time token-by-token responses to hold up as usage grew, alongside getting markdown rendering fully correct in the chat UI.



Impact
The result reads less like a search box and more like a conversation — answers stream in immediately and stay grounded in Inspigo's own learning content rather than drifting into generic, ungrounded responses. Built with a 5-person team; peak usage reaches thousands of conversations per day.