About MokingBird AI
Local AI Infrastructure, Built for the Real World
MokingBird AI is the desktop AI ecosystem developed by MokingBird — a suite of three powerful, production-grade tools for working with large language models, entirely on your own hardware. No cloud accounts. No data leaving your machine. No subscriptions required to get started.
The suite is organized around the MokingBird Node — a desktop hub application that brings mbRAG, mbDataGen, and mbFT together in one place. Use them as an integrated system through the Node, or run each tool independently as a standalone application.
Domain: ai.mokingbird.xyz Company: MokingBird Oy, Business ID: 3615646-1, Finland
Our Mission
Your models. Your hardware. Your data. Your rules.
Local-first AI infrastructure for builders, researchers, and teams who need professional capability without cloud lock-in or data exposure.
What is MokingBird AI?
MokingBird AI is an umbrella for our local AI infrastructure products. The core architecture is:
MokingBird Node (Desktop Hub)
├── mbRAG — Retrieval-Augmented Generation framework
├── mbDataGen — Synthetic dataset generation platform
└── mbFT — Universal fine-tuning platform
Each tool can be installed and used standalone, or together through the Node. The Node is the hub — a PySide6 desktop application (Python) that surfaces all three tools in a unified interface and exposes local FastAPI endpoints for programmatic access.
Platforms: Windows, macOS, Linux Architecture: 100% local execution — all AI operations run on your machine Internet: Optional (for connecting to cloud LLM providers using your own API keys)
Our Three Tools
mbRAG — Advanced Retrieval-Augmented Generation
mbRAG is a production-ready, stable alternative to LangChain. It implements every major RAG approach in a single unified system — without the breaking changes, vendor lock-in, or opacity that plagues other frameworks.
At a glance:
- 4 pipeline levels: L1 Basic (~0.8s) → L4 Advanced (~3.5s)
- 17 supported document formats (PDF with 3 engine fallbacks, DOCX, Excel, CSV, JSON, Markdown, PowerPoint, Images with OCR, Web content, and more)
- 10 LLM providers (OpenAI, Anthropic Claude, Google Gemini, Ollama, HuggingFace, vLLM, llama.cpp, and more)
- 8 embedding providers
- 6 vector store backends (ChromaDB, FAISS, Qdrant, Pinecone, Weaviate, Milvus)
- 6 retrieval strategies including Ensemble and Contextual
- 6-Level Contextual Retrieval — MokingBird's signature innovation that enriches every document chunk with surrounding context before retrieval, dramatically improving accuracy
- 40–50% accuracy improvement over naive RAG implementations
mbDataGen — Synthetic Dataset Generation
mbDataGen solves one of the most persistent problems in applied ML: getting high-quality, domain-specific training data. Instead of scraping, labeling manually, or accepting noisy public datasets, mbDataGen generates clean, validated synthetic data from your own documents.
At a glance:
- 5-phase pipeline: Extract → Enrich → Generate → Validate → Deploy
- GPRO-Hybrid RL — our original reward learning approach:
Total Reward = 0.7 × Field/Process Reward + 0.3 × Outcome/Overall Reward - K=4 candidate generation with comparative reward scoring
- 5-stage validator: Schema → Distribution → Dedupe → Grounding → Novelty
- HMAC-signed RunManifest for complete data provenance
- Minimum 6GB VRAM, 8GB recommended
- Generates output in any schema you define — Jogg quiz format, instruction-following pairs, preference data, and more
mbFT — Universal Fine-Tuning Platform
mbFT makes fine-tuning accessible without hiding what's actually happening. The Smart Config Engine handles the complexity of choosing hyperparameters and memory-efficient configurations, while keeping you in full control.
At a glance:
- 16 fine-tuning techniques:
- 6 SFT methods (LoRA, QLoRA, Full Fine-Tuning, Prefix Tuning, Prompt Tuning, Adapter Layers)
- 5 RL methods (GRPO, PPO, DPO, ORPO, Kahneman-Tversky Optimization)
- 5 Multimodal methods (Vision-Language, Audio-Text, Code-Specialized, Medical Imaging, Document Understanding)
- 7 supported frameworks (Unsloth, Axolotl, LLaMAFactory, Transformers, DeepSpeed, FSDP, TRL)
- VRAM pre-simulation — before you start a run, the system estimates memory requirements so you know if your hardware can handle it
- Hybrid GRPO — MokingBird's original contribution, combining reward model and rule-based signals
- 6-tier hardware classification system
- Comparison interface vs manual setups in Unsloth/Axolotl/LLaMAFactory
Why Local AI?
The cloud AI model has a hidden cost: your data. When you send documents to an API endpoint to answer questions, generate datasets, or fine-tune a model, those documents leave your control. Even with privacy-protecting terms of service, the fundamental architecture means your data travels.
MokingBird AI is built on a different premise: the model comes to your data, not the other way around.
- Your documents stay on your machine
- Your API keys are stored locally, never transmitted to MokingBird
- Your fine-tuned models are yours — we have no access to them
- No registration required
- Works fully offline when using local LLMs via Ollama
This isn't just a privacy positioning. It's an architecture choice that also means lower latency, no per-query costs for local models, and no dependency on service uptime. Critical workflows don't become unusable when an external API is down.
Our Mission for AI
We believe powerful AI tools should not require:
- A cloud account
- A corporate API budget
- Trusting a third party with proprietary data
- A PhD to configure
MokingBird AI is our contribution toward democratizing serious AI infrastructure — the kind of RAG, data generation, and fine-tuning capability that has historically been available only to well-funded research teams or large enterprises.
Part of MokingBird
MokingBird AI is developed by MokingBird Oy — The Everything Lab. We also build MB Viewer, Sortify, Jogg, and Jogg Mini.
Learn more about MokingBird at mokingbird.xyz/about.
Contact
- General: [email protected]
- Support: [email protected]
- Enterprise inquiries: [email protected]
- Company-level: [email protected]