Back to Projects
Flarif AI
AI/ML

Flarif AI

Senior AI/Software Engineer & Tech Lead

Next.jsNestJSLangChainOpenAIRAGKnowledge GraphsSocket.io

About the Project

Flarif AI lets individuals and businesses build secure, custom-trained AI assistants over their own documents. What began as a RAG product grew into an agentic platform: an orchestration layer decides per turn whether a question needs retrieval, a live web search, a multi-round research pass, or nothing but the conversation so far. The product spans eight services — a socket server running the AI orchestration engine, an Express/MongoDB backend, a NestJS platform API, the Next.js web app, an admin console, a masterclass site, a HubSpot app, and a WordPress plugin — with distinct conversational surfaces for personal use, projects, business widgets, and deployments.

Key Highlights

  • Built the orchestration layer: an LLM-driven classifier decides per turn whether to retrieve, search the web, or run a bounded multi-round deep-research pass with a clarification step before spending tokens
  • Designed and implemented the full RAG pipeline — ingestion, chunking, embedding, vector storage, retrieval — alongside a knowledge-graph service for relationship-aware answers
  • Implemented context compaction so long conversations stay inside the window, deliberately routed to a cheap fixed model independent of whichever model the conversation itself is using
  • Built multi-model routing across providers via a model catalogue, plus a request extension that lets a downstream provider skip its own duplicate classification when this layer has already made the call
  • Shipped distribution beyond the web app — HubSpot app, WordPress plugin, embeddable business widgets, and Slack, Telegram, Discord, and Meta integrations — with human escalation when the agent should hand off

Technical Challenges

The hard problem stopped being retrieval quality and became knowing when to retrieve at all. Running search and deep research on every turn is slow and expensive; running them on none makes the assistant useless on anything current. So the decision itself became a classified step, with bounded rounds so a research pass terminates instead of spiralling. The cost discipline that follows from that is the part I'd defend in review: background work — compaction, classification, titling — is pinned to a cheap model regardless of what the user's conversation is running on, because the user never sees that output and paying frontier prices for bookkeeping is money set on fire.