Saepul Malik
Case study

Inspigo AI: Real-Time RAG for Personalized Learning

A streaming chatbot experience powered by OpenAI/Bedrock and a RAG pipeline, grounded in the platform’s own learning content.

PT Inspigo Inovasi Indonesia2021–Present
Next.jsOpenAIBedrockRAG

Problem

Personalized learning benefits from conversational, context-aware assistance rather than static content lookup — but a foundation model answering purely from its general training data isn't grounded in what the platform actually teaches.

Architecture

Source documents are uploaded through the Inspigo CMS, stored in S3, and ingested into Pinecone as the vector store — Pinecone handles chunking and metadata, with embeddings generated via OpenAI's embedding models. At generation time, retrieved chunks are injected into the prompt before the model responds. OpenAI is the primary model, with AWS Bedrock as a fallback.

Execution

Owned the frontend/streaming UX end to end, plus the CMS interface used to upload and manage source documents. Streaming shipped first over WebSocket (still powering some legacy roleplay features), with newer flows moved to SSE. The main hard problem was scaling the streaming layer under concurrent load — getting real-time token-by-token responses to hold up as usage grew, alongside getting markdown rendering fully correct in the chat UI.

Inspigo AI chat conversation

Inspigo AI roleplay detail

Inspigo AI evaluation

Impact

The result reads less like a search box and more like a conversation — answers stream in immediately and stay grounded in Inspigo's own learning content rather than drifting into generic, ungrounded responses. Built with a 5-person team; peak usage reaches thousands of conversations per day.