Configuration

Sirchmunk is configured through environment variables stored in a .env file. After running sirchmunk init, the configuration file is created at ~/.sirchmunk/.env.

Environment Variables

LLM Configuration

VariableDescriptionDefault
LLM_API_KEYYour LLM API key (required for FAST and DEEP modes)—
LLM_BASE_URLOpenAI-compatible API base URLhttps://api.openai.com/v1
LLM_MODEL_NAMEModel name to usegpt-5.2

Search Configuration

VariableDescriptionDefault
SIRCHMUNK_WORK_PATHWorking directory for data storage~/.sirchmunk/
SIRCHMUNK_SEARCH_PATHSDefault search paths (comma-separated)—
SIRCHMUNK_MAX_DEPTHMaximum directory traversal depth10
SIRCHMUNK_TOP_K_FILESNumber of top files to analyze20
SIRCHMUNK_MAX_CONCURRENT_SEARCHESMax concurrent search tasks3
SIRCHMUNK_ENABLE_CLUSTER_REUSEEnable knowledge cluster reusetrue

Retrieval Cost Configuration

VariableDescriptionDefault
GREP_MAX_FILESIZE_MBPer-file size cap (MB); files over this limit are skipped on the query hot path64
GREP_RGA_ADAPTERSAllowed rga adapters (bounded only; archives disabled on hot path)poppler,pandoc,postprocpagebreaks
GREP_TIERED_SCANEnable tiered scan: fast native-rg pass + bounded rga pass for rich formatstrue
GREP_TEXT_TIMEOUTTimeout (seconds) for the native rg text pass15.0
GREP_TIMEOUTTimeout (seconds) for the rga rich pass60.0
GREP_RICH_EXTENSIONSFile extensions routed to the rga rich passpdf,docx,epub,odt

Chat Configuration

VariableDescriptionDefault
CHAT_HISTORY_MAX_TURNSMaximum number of chat turns retained in history—
CHAT_HISTORY_MAX_TOKENSMaximum token budget for retained chat history—

Server Configuration

VariableDescriptionDefault
SIRCHMUNK_HOSTAPI server bind address127.0.0.1
SIRCHMUNK_PORTAPI server port8584

Data Storage Layout

All persistent data is stored under SIRCHMUNK_WORK_PATH:

{SIRCHMUNK_WORK_PATH}/
  ├── .cache/
  │   ├── history/              # Chat session history (DuckDB)
  │   │   └── chat_history.db
  │   ├── knowledge/            # Knowledge clusters (Parquet)
  │   │   └── knowledge_clusters.parquet
  │   ├── compile/              # Compile artifacts (Beta)
  │   │   ├── manifest.json     # File manifest with hashes
  │   │   ├── document_catalog.json
  │   │   ├── summary_index.json
  │   │   ├── trees/            # Hierarchical tree indices
  │   │   ├── table_digests/    # Table extraction digests
  │   │   └── xlsx_digests/     # Spreadsheet digests
  │   └── settings/             # User settings (DuckDB)
  │       └── settings.db
  ├── .env                      # Environment configuration
  └── mcp_config.json           # MCP server configuration

Search Parameters

When invoking search (via SDK, CLI, or API), the following parameters are available:

ParameterTypeDefaultDescription
querystringrequiredSearch query or question
pathsstring | string[]optionalDirectories or files to search; falls back to SIRCHMUNK_SEARCH_PATHS, then cwd
modestringDEEPDEEP (agentic retrieval with budgeted evidence exploration), FAST (greedy, 2-5s), or FILENAME_ONLY
max_depthintnullMaximum directory depth
top_k_filesintnullNumber of top files to return
enable_dir_scanbooltrueEnable directory scanning
max_loopsintnullDEEP mode loop limit
max_token_budgetintnullDEEP mode token budget (default 128K when unset)
include_patternsstring[]nullFile glob patterns to include
exclude_patternsstring[]nullFile glob patterns to exclude
response_formatstringrich"rich" Markdown report, "minimal" short answer, "context" SearchContext, or "json" serialized context
Note

FILENAME_ONLY mode does not require an LLM API key. FAST and DEEP modes require a configured LLM. Default mode is DEEP, which performs budgeted evidence exploration with multi-path retrieval.

docs