FAQ
Frequently asked questions about Sirchmunk.
How is this different from traditional RAG systems?
Sirchmunk takes an indexless approach:
- No pre-indexing: Direct file search without vector database setup
- Self-evolving: Knowledge clusters evolve based on search patterns
- Multi-level retrieval: Adaptive keyword granularity for better recall
- Evidence-based: Budgeted evidence localization for precise, source-grounded extraction
What LLM providers are supported?
Any OpenAI-compatible API endpoint, including:
- OpenAI (GPT-4, GPT-4o, GPT-5.2)
- Local models served via Ollama, llama.cpp, vLLM, SGLang
- Claude via API proxy
- MiniMax (MiniMax-M3, MiniMax-M2.7, MiniMax-M2.7-highspeed)
- DeepSeek, Moonshot, Mistral, Groq, Together AI, Cohere
- Google Gemini, Zhipu (GLM), Baichuan, Yi, SiliconFlow, Volcengine, Azure OpenAI
- Any other OpenAI-compatible provider
How do I add documents to search?
Simply specify the path in your search query — no pre-processing or indexing required:
result = await searcher.search(
query="Your question",
paths=["/path/to/folder", "/path/to/file.pdf"]
)
Where are knowledge clusters stored?
Knowledge clusters are persisted in Parquet format at:
{SIRCHMUNK_WORK_PATH}/.cache/knowledge/knowledge_clusters.parquet
You can query them using DuckDB or the KnowledgeManager API.
How do I monitor LLM token usage?
Three ways:
- Web Dashboard: Visit the Monitor page for real-time statistics
- API:
GET /api/v1/monitor/llmreturns usage metrics - Code: Access
searcher.llm_usagesafter search completion
Does FILENAME_ONLY mode require an LLM?
No. FILENAME_ONLY mode performs fast filename-based search without any LLM calls. Both FAST and DEEP modes require a configured LLM API key. Default mode is now DEEP, which performs budgeted evidence exploration with multi-path retrieval.
What file formats are supported?
Sirchmunk leverages ripgrep-all to search across 100+ file formats, including:
- Source code (Python, JavaScript, Java, Go, Rust, etc.)
- Documents (PDF, DOCX, XLSX, PPTX)
- Archives (ZIP, TAR, GZ)
- Data files (JSON, YAML, CSV, XML)
- Plain text (TXT, MD, RST)
- And many more
How does the knowledge system evolve?
Knowledge clusters follow a natural lifecycle:
- Creation — New evidence generates a new cluster
- Reuse — Similar queries match and enhance existing clusters
- Maturation — Repeated validation transitions clusters from Emerging to Stable
- Deprecation — Unsupported clusters transition to Contested or Deprecated
What is the LENS paper and how does it relate to Sirchmunk?
LENS (Latent Evidence Exploration and Search) is the research paper that formalizes Sirchmunk’s core retrieval algorithm. Published as arXiv:2608.16185, it describes how Sirchmunk performs budgeted evidence localization over dynamic raw documents without pre-built indexes. Read the full paper at https://arxiv.org/abs/2608.16185.