Actions
37 actions, callable from the panel, the command palette and the API as
POST /api/v1/a/<id>. Internal actions used between modules are not listed.
| Action | What it does | Risk | Preview |
|---|---|---|---|
ai.status | AI status and setup state (read) | low | No dry run |
ai.policy.get | Get the AI policy (read) | low | No dry run |
ai.policy.set | Change the AI policy | high | No dry run |
ai.killswitch.set | Turn the assistant or its changes on or off | high | No dry run |
ai.provider.list | List model providers (read) | low | No dry run |
ai.provider.create | Add a model provider | high | No dry run |
ai.provider.update | Change a model provider | high | No dry run |
ai.provider.delete | Remove a model provider | high | No dry run |
ai.provider.test | Test a provider (list models and a tiny completion) (read) | low | No dry run |
ai.model.list | List catalogued models and routes (read) | low | No dry run |
ai.model.configure | Add or change a catalogued model | medium | No dry run |
ai.model.delete | Remove a catalogued model | medium | No dry run |
ai.route.set | Choose the models a feature uses | medium | No dry run |
ai.chat.send | Send a message to the assistant | low | No dry run |
ai.turn.get | Read an assistant turn (read) | low | No dry run |
ai.tools.list | List the tools your assistant may use (read) | low | No dry run |
ai.session.list | List your conversations (read) | low | No dry run |
ai.session.get | Read (export) a conversation (read) | low | No dry run |
ai.session.delete | Delete a conversation | low | No dry run |
ai.plan.get | Get a plan with its diffs (read) | low | No dry run |
ai.plan.approve | Approve and run a plan | medium | No dry run |
ai.plan.reject | Reject a plan | low | No dry run |
ai.plan.undo | Undo an executed plan | medium | No dry run |
ai.rag.search | Search the documentation (read) | low | No dry run |
ai.rag.status | Documentation index status (read) | low | No dry run |
ai.rag.reindex | Rebuild the documentation index | medium | No dry run |
ai.kb.create | Create a knowledge base | medium | No dry run |
ai.kb.list | List knowledge bases (read) | low | No dry run |
ai.kb.delete | Delete a knowledge base and all its data | medium | No dry run |
ai.kb.doc.add | Add a document to a knowledge base | low | No dry run |
ai.kb.doc.delete | Remove a document and its vectors | low | No dry run |
ai.kb.query | Ask a knowledge base (RAG) (read) | low | No dry run |
ai.gateway.key.create | Create an AI gateway key | medium | No dry run |
ai.gateway.key.list | List AI gateway keys (read) | low | No dry run |
ai.gateway.key.revoke | Revoke an AI gateway key | medium | No dry run |
ai.gateway.logs | AI gateway request log (read) | low | No dry run |
ai.usage.get | AI usage and credits (read) | low | No dry run |
Permissions and limits
Permissions
ai.assistant.useUse the assistantai.plan.approveApprove and undo assistant plansai.provider.manageManage model providers and routingai.policy.manageManage the AI policy and kill switchai.rag.adminManage the documentation indexai.usage.viewSee AI usage and creditsai.gateway.manageManage AI gateway keysai.kb.manageManage knowledge basesai.kb.useQuery knowledge bases
Plan limits
ai.credits_monthAI credits per monthai.assistant_msgs_dayAssistant messages per dayai.gateway_keysAI gateway keysai.kb_countKnowledge bases
Engineering notes
Generated from modules/ai/docs.md at build d90e9e2. These are the notes the engineers keep
next to the code: precise, technical, and honest about what is not done yet.
Ops assistant, documentation RAG, model providers, customer AI gateway and knowledge bases. Threat model:
docs/security/ai.md. Card: docs/tasks/w8/w8-03-ai.md. Edition core (works on every licence; customer services are
gated by the plan's ai.credits_month, default 0).
Layout
llm/model layer (adapted from nerixchat/llm.py,common/providers.py):Model/Embedderinterfaces, oneClientfor Anthropic Messages (prompt caching: system + last tool + turns[-2] breakpoints; usage reconciled to four slots) and the OpenAI wire (OpenAI, DeepSeek, Mistral, Gemini OpenAI endpoint, OpenRouter, any compatible URL, Ollama/vLLM/llama.cpp).Egress: retries 429/5xx with Retry-After, provider up/down board, body guard.safety/redaction, egress guard, injection scan, untrusted-data fence.rag/heading-aware chunking (1,800/200/200 chars, heading path carried forward), built-in feature-hash embedder (builtin-hash-256, offline), reciprocal-rank fusion (k=60). Adapted from nerixlibrary/*.assistant/tools from the registry, the turn loop (quarantine), plan hash, step-wise execution with re-preview.cp/the plugin: actions, Postgres store (schemam_ai), pgvector or exact vector spaces, gateway listener.mockmodel/deterministic OpenAI-compatible mock (lab containerrc-w803-mock; a deliberately naive "model").tests/lab.shlab proofs (up | run | down).
How it works
- Turn:
ai.chat.send {message, session_id?, context?, wait?}-> tools = registry actions the caller may run (ai.tools.listshows them and why others were skipped) -> model loop (<= 6 calls). Reads run as the caller; a mutating call runs the action's **dry run** and becomes a plan step. Live progress: eventai.turn.updated(ids only) ->ai.turn.get {id, after_seq}returns new parts + the streamingdraft. - Plan:
ai.plan.get(steps, params, previews, risk, expiry, reversible) ->ai.plan.approve {id, hash}(step-up for high/critical) -> each step re-previewed (must match), executed, verified (<module>.getwhen it exists), inverse recorded ->ai.plan.undo {id}withinundo_days. - RAG: corpus = every installed module's
docs.md+ Markdown under/usr/share/respirecloud/docs(settingdocs_dirs). Keyword pane (Postgres FTS, >= half the query terms) + semantic pane (active space) fused by rank; no relevant passage -> abstain. A space is one embedding model; changing the routed model builds a new space beside the active one and swaps when complete. pgvector (halfvec(dims)+ HNSW m=16/ef 64, iterative scans) when the extension exists inm_ai, exact cosine otherwise. Refreshed every 30 min (ai.docs.refresh) and onai.rag.reindex. - Gateway:
POST /v1/chat/completions(stream or not),/v1/embeddings,GET /v1/models(also under/ai/v1) on127.0.0.1:7790; keyrcai_<id>_<secret>; modelauto= thegatewayroute. Usage read from the provider's final chunk (stream_options.include_usageforced) and metered. - Metering: every call ->
usage_ledger(4 token slots, credits = rate table per 1M tokens; estimates). Non-admins are capped byai.credits_month(assistant, KB answers, gateway) andai.assistant_msgs_day.
Verified (lab w8-03, 2026-10-10)
Docs question answered with citations (docs/specs/ai.md › … › Assistant: safety model); out-of-corpus question
abstains; "add domain shop-alice.lab.test for alice" -> plan with the domains dry-run diff -> wrong hash refused -> approved
-> executed, verified, audited (ai.plan.proposed/approved/executed, domains.create by admin) -> re-approve refused ->
undo removed the domain; injection line in alice's log -> quarantined, naive domains_delete blocked, no plan, audit
ai.injection.suspected (denied), 0 secrets at the model server; gateway key for alice: 429 insufficient_quota with
plan credits 0, after assigning 1 credit: 200 + streamed 200, then 429 with Retry-After, ledger gateway 1.235;
kill switch refuses chat; pgvector space built and swapped in when the extension exists.
Next
UI (assistant panel, plan review, providers, usage); data-egress consent per category; local model runtime
(ai.local.* agent ops); managed vector collections; per-server config/log corpus via agent ops; reseller-scoped keys.