AI assistant and services

Ops assistant (plan, diff, approve), docs search, model providers and a per-account OpenAI-compatible gateway.

In preview ai version 0.1.0

Actions

37 actions, callable from the panel, the command palette and the API as POST /api/v1/a/<id>. Internal actions used between modules are not listed.

ActionWhat it doesRiskPreview
ai.status AI status and setup state (read) low No dry run
ai.policy.get Get the AI policy (read) low No dry run
ai.policy.set Change the AI policy high No dry run
ai.killswitch.set Turn the assistant or its changes on or off high No dry run
ai.provider.list List model providers (read) low No dry run
ai.provider.create Add a model provider high No dry run
ai.provider.update Change a model provider high No dry run
ai.provider.delete Remove a model provider high No dry run
ai.provider.test Test a provider (list models and a tiny completion) (read) low No dry run
ai.model.list List catalogued models and routes (read) low No dry run
ai.model.configure Add or change a catalogued model medium No dry run
ai.model.delete Remove a catalogued model medium No dry run
ai.route.set Choose the models a feature uses medium No dry run
ai.chat.send Send a message to the assistant low No dry run
ai.turn.get Read an assistant turn (read) low No dry run
ai.tools.list List the tools your assistant may use (read) low No dry run
ai.session.list List your conversations (read) low No dry run
ai.session.get Read (export) a conversation (read) low No dry run
ai.session.delete Delete a conversation low No dry run
ai.plan.get Get a plan with its diffs (read) low No dry run
ai.plan.approve Approve and run a plan medium No dry run
ai.plan.reject Reject a plan low No dry run
ai.plan.undo Undo an executed plan medium No dry run
ai.rag.search Search the documentation (read) low No dry run
ai.rag.status Documentation index status (read) low No dry run
ai.rag.reindex Rebuild the documentation index medium No dry run
ai.kb.create Create a knowledge base medium No dry run
ai.kb.list List knowledge bases (read) low No dry run
ai.kb.delete Delete a knowledge base and all its data medium No dry run
ai.kb.doc.add Add a document to a knowledge base low No dry run
ai.kb.doc.delete Remove a document and its vectors low No dry run
ai.kb.query Ask a knowledge base (RAG) (read) low No dry run
ai.gateway.key.create Create an AI gateway key medium No dry run
ai.gateway.key.list List AI gateway keys (read) low No dry run
ai.gateway.key.revoke Revoke an AI gateway key medium No dry run
ai.gateway.logs AI gateway request log (read) low No dry run
ai.usage.get AI usage and credits (read) low No dry run

Permissions and limits

Permissions

  • ai.assistant.use Use the assistant
  • ai.plan.approve Approve and undo assistant plans
  • ai.provider.manage Manage model providers and routing
  • ai.policy.manage Manage the AI policy and kill switch
  • ai.rag.admin Manage the documentation index
  • ai.usage.view See AI usage and credits
  • ai.gateway.manage Manage AI gateway keys
  • ai.kb.manage Manage knowledge bases
  • ai.kb.use Query knowledge bases

Plan limits

  • ai.credits_month AI credits per month
  • ai.assistant_msgs_day Assistant messages per day
  • ai.gateway_keys AI gateway keys
  • ai.kb_count Knowledge bases

Engineering notes

Generated from modules/ai/docs.md at build d90e9e2. These are the notes the engineers keep next to the code: precise, technical, and honest about what is not done yet.

Ops assistant, documentation RAG, model providers, customer AI gateway and knowledge bases. Threat model: docs/security/ai.md. Card: docs/tasks/w8/w8-03-ai.md. Edition core (works on every licence; customer services are gated by the plan's ai.credits_month, default 0).

Layout

  • llm/ model layer (adapted from nerix chat/llm.py, common/providers.py): Model/Embedder interfaces, one Client for Anthropic Messages (prompt caching: system + last tool + turns[-2] breakpoints; usage reconciled to four slots) and the OpenAI wire (OpenAI, DeepSeek, Mistral, Gemini OpenAI endpoint, OpenRouter, any compatible URL, Ollama/vLLM/llama.cpp). Egress: retries 429/5xx with Retry-After, provider up/down board, body guard.
  • safety/ redaction, egress guard, injection scan, untrusted-data fence.
  • rag/ heading-aware chunking (1,800/200/200 chars, heading path carried forward), built-in feature-hash embedder (builtin-hash-256, offline), reciprocal-rank fusion (k=60). Adapted from nerix library/*.
  • assistant/ tools from the registry, the turn loop (quarantine), plan hash, step-wise execution with re-preview.
  • cp/ the plugin: actions, Postgres store (schema m_ai), pgvector or exact vector spaces, gateway listener.
  • mockmodel/ deterministic OpenAI-compatible mock (lab container rc-w803-mock; a deliberately naive "model").
  • tests/lab.sh lab proofs (up | run | down).

How it works

  • Turn: ai.chat.send {message, session_id?, context?, wait?} -> tools = registry actions the caller may run (ai.tools.list shows them and why others were skipped) -> model loop (<= 6 calls). Reads run as the caller; a mutating call runs the action's **dry run** and becomes a plan step. Live progress: event ai.turn.updated (ids only) -> ai.turn.get {id, after_seq} returns new parts + the streaming draft.
  • Plan: ai.plan.get (steps, params, previews, risk, expiry, reversible) -> ai.plan.approve {id, hash} (step-up for high/critical) -> each step re-previewed (must match), executed, verified (<module>.get when it exists), inverse recorded -> ai.plan.undo {id} within undo_days.
  • RAG: corpus = every installed module's docs.md + Markdown under /usr/share/respirecloud/docs (setting docs_dirs). Keyword pane (Postgres FTS, >= half the query terms) + semantic pane (active space) fused by rank; no relevant passage -> abstain. A space is one embedding model; changing the routed model builds a new space beside the active one and swaps when complete. pgvector (halfvec(dims) + HNSW m=16/ef 64, iterative scans) when the extension exists in m_ai, exact cosine otherwise. Refreshed every 30 min (ai.docs.refresh) and on ai.rag.reindex.
  • Gateway: POST /v1/chat/completions (stream or not), /v1/embeddings, GET /v1/models (also under /ai/v1) on 127.0.0.1:7790; key rcai_<id>_<secret>; model auto = the gateway route. Usage read from the provider's final chunk (stream_options.include_usage forced) and metered.
  • Metering: every call -> usage_ledger (4 token slots, credits = rate table per 1M tokens; estimates). Non-admins are capped by ai.credits_month (assistant, KB answers, gateway) and ai.assistant_msgs_day.

Verified (lab w8-03, 2026-10-10)

Docs question answered with citations (docs/specs/ai.md › … › Assistant: safety model); out-of-corpus question abstains; "add domain shop-alice.lab.test for alice" -> plan with the domains dry-run diff -> wrong hash refused -> approved -> executed, verified, audited (ai.plan.proposed/approved/executed, domains.create by admin) -> re-approve refused -> undo removed the domain; injection line in alice's log -> quarantined, naive domains_delete blocked, no plan, audit ai.injection.suspected (denied), 0 secrets at the model server; gateway key for alice: 429 insufficient_quota with plan credits 0, after assigning 1 credit: 200 + streamed 200, then 429 with Retry-After, ledger gateway 1.235; kill switch refuses chat; pgvector space built and swapped in when the extension exists.

Next

UI (assistant panel, plan review, providers, usage); data-egress consent per category; local model runtime (ai.local.* agent ops); managed vector collections; per-server config/log corpus via agent ops; reseller-scoped keys.