ikramdeveloper

Ikramdeveloper

All Projects

Next.js, NestJS, Python, Qdrant, RAG, Paddle

Ultranivo — Branded AI Chatbot SaaS

Ultranivo is a multi-tenant SaaS I founded and built solo: organizations add documents, web pages, and YouTube videos, and get a branded AI assistant — on their own domain or embedded on their site — that answers their audience's questions with citations to the exact page of a document or the exact timestamp of a video.

Ultranivo
The Problem

What needed to be solved

General AI tools answer from the whole internet, under someone else's brand. Businesses, non-profits, academies, creators — anyone with knowledge or information to share, or earn from — had no simple way to give their audience an AI experience grounded only in their own content, under their own brand, with the option to charge for access. Building that requires a full RAG pipeline, multi-tenant isolation, ingestion for multiple content formats, and billing — far beyond a no-code audience.

Engineering Thinking

Key Decisions & Challenges

1

Database-per-tenant isolation

Situation

Each customer's content and chats must be strictly isolated — a tenant's AI must never answer from another tenant's data — while global config (plans, models, tenants) stays shared.

Options Considered

  • Shared collections with a tenantId field on every document
  • One dynamically-connected MongoDB database per tenant

Decision

One MongoDB database per tenant, connected dynamically at request time from the JWT claim or admin query param, plus per-tenant Qdrant vector collections derived from the tenant's immutable database name. A static 'super DB' holds cross-tenant models. Isolation is structural rather than query-discipline, and a tenant can be exported or dropped wholesale.

2

No database access in the ingestion services

Situation

Four independently-deployed Python services (embedding, video, document processing) originally needed persistence, but sharing a database across two codebases meant duplicated schemas and drift.

Options Considered

  • Direct database access from Python with duplicated models
  • A dedicated ingestion REST API on the NestJS backend

Decision

Removed all database access from Python. Workers discover queued jobs and report progress, status, and chunks over an authenticated HTTP API owned by the backend — one source of truth for schemas and contracts, and the services stay stateless and independently deployable.

3

Fallback chains for every AI dependency

Situation

The platform depends on third-party AI providers for chat, OCR, and transcription — any single provider's outage or rate limit would take the product down.

Decision

Every AI call runs through a fallback chain: chat answers via Gemini with OpenRouter and Groq fallbacks; document OCR via Gemini Vision, then OpenRouter, then local EasyOCR; video transcripts from YouTube captions with a Whisper fallback. Provider/model config is managed live from the admin panel, with per-provider health monitoring.

What Was Built

Features

Cited RAG Chat

Semantic search over the tenant's content (bge-m3 embeddings + Qdrant + reranking) feeds the LLM, and every answer cites its sources — down to the exact page of a document or timestamp link in a video.

Multi-Format Ingestion Pipeline

PDFs and DOCX through an OCR pipeline, pasted text, web pages, and YouTube videos via caption extraction with Whisper transcription fallback — all queued, processed by workers, and tracked with live progress.

Embeddable Widget

A drop-in script tag that embeds the tenant's branded chat assistant on any existing website.

White-Label Tenancy

Each customer gets their own branded assistant on their own domain, with per-tenant databases and vector collections keeping content fully isolated.

Self-Serve Billing

Signup → plan selection → Paddle checkout, with subscription plans, coupons, credit rates, and usage tracking managed from the admin panel.

Admin & Monitoring Dashboard

Internal dashboard for content, users, chats, AI model configuration, ingestion workers, cost estimation, and health monitoring of the vector database and AI providers.

Internationalization

The chat product ships in English and Urdu.

Screenshots

In Action

Platform overview

Platform overview

Setting up an assistant

Setting up an assistant

Adding content

Adding content

Monetization

Monetization

Responsibility

My Role

Founder and solo engineer. Designed, built, launched, and operate the entire platform — two Next.js apps (tenant chat UI and admin panel + marketing site), a NestJS API, and four Python ingestion services — plus billing, infrastructure, and go-to-market.

Stack

Tech Stack

Next.js React TypeScript TailwindCSS Shadcn Zustand React Query NestJS Node.js MongoDB Qdrant Python FastAPI Redis Paddle AWS S3 Resend
Result

Outcome

Live in production at ultranivo.com — self-serve signup with paid monthly and yearly plans.

Demo

Live Project

www.ultranivo.com