Next.js, NestJS, Python, Qdrant, RAG, Paddle
Ultranivo — Branded AI Chatbot SaaS
Ultranivo is a multi-tenant SaaS I founded and built solo: organizations add documents, web pages, and YouTube videos, and get a branded AI assistant — on their own domain or embedded on their site — that answers their audience's questions with citations to the exact page of a document or the exact timestamp of a video.
Resources
What needed to be solved
General AI tools answer from the whole internet, under someone else's brand. Businesses, non-profits, academies, creators — anyone with knowledge or information to share, or earn from — had no simple way to give their audience an AI experience grounded only in their own content, under their own brand, with the option to charge for access. Building that requires a full RAG pipeline, multi-tenant isolation, ingestion for multiple content formats, and billing — far beyond a no-code audience.
Key Decisions & Challenges
Database-per-tenant isolation
Situation
Each customer's content and chats must be strictly isolated — a tenant's AI must never answer from another tenant's data — while global config (plans, models, tenants) stays shared.
Options Considered
- Shared collections with a tenantId field on every document
- One dynamically-connected MongoDB database per tenant
Decision
One MongoDB database per tenant, connected dynamically at request time from the JWT claim or admin query param, plus per-tenant Qdrant vector collections derived from the tenant's immutable database name. A static 'super DB' holds cross-tenant models. Isolation is structural rather than query-discipline, and a tenant can be exported or dropped wholesale.
No database access in the ingestion services
Situation
Four independently-deployed Python services (embedding, video, document processing) originally needed persistence, but sharing a database across two codebases meant duplicated schemas and drift.
Options Considered
- Direct database access from Python with duplicated models
- A dedicated ingestion REST API on the NestJS backend
Decision
Removed all database access from Python. Workers discover queued jobs and report progress, status, and chunks over an authenticated HTTP API owned by the backend — one source of truth for schemas and contracts, and the services stay stateless and independently deployable.
Fallback chains for every AI dependency
Situation
The platform depends on third-party AI providers for chat, OCR, and transcription — any single provider's outage or rate limit would take the product down.
Decision
Every AI call runs through a fallback chain: chat answers via Gemini with OpenRouter and Groq fallbacks; document OCR via Gemini Vision, then OpenRouter, then local EasyOCR; video transcripts from YouTube captions with a Whisper fallback. Provider/model config is managed live from the admin panel, with per-provider health monitoring.
Features
Cited RAG Chat
Semantic search over the tenant's content (bge-m3 embeddings + Qdrant + reranking) feeds the LLM, and every answer cites its sources — down to the exact page of a document or timestamp link in a video.
Multi-Format Ingestion Pipeline
PDFs and DOCX through an OCR pipeline, pasted text, web pages, and YouTube videos via caption extraction with Whisper transcription fallback — all queued, processed by workers, and tracked with live progress.
Embeddable Widget
A drop-in script tag that embeds the tenant's branded chat assistant on any existing website.
White-Label Tenancy
Each customer gets their own branded assistant on their own domain, with per-tenant databases and vector collections keeping content fully isolated.
Self-Serve Billing
Signup → plan selection → Paddle checkout, with subscription plans, coupons, credit rates, and usage tracking managed from the admin panel.
Admin & Monitoring Dashboard
Internal dashboard for content, users, chats, AI model configuration, ingestion workers, cost estimation, and health monitoring of the vector database and AI providers.
Internationalization
The chat product ships in English and Urdu.
In Action
Platform overview
Setting up an assistant
Adding content
Monetization
My Role
Founder and solo engineer. Designed, built, launched, and operate the entire platform — two Next.js apps (tenant chat UI and admin panel + marketing site), a NestJS API, and four Python ingestion services — plus billing, infrastructure, and go-to-market.
Tech Stack
Outcome
Live in production at ultranivo.com — self-serve signup with paid monthly and yearly plans.