11 KiB
AI Agent App — product specification
This is the contract the project is built against. It is written for the agent that scaffolds the repository, but it doubles as the human-readable brief: every requirement below is meant to be implementable and verifiable.
Values in angle brackets (<Agent name>, <model>, <retrieval>) come from the
project-creation dialog and are recorded in the project record; replace them as you read.
1. Goal
A chat application built around an LLM agent: a user holds a conversation, the agent streams its
answer token by token, may consult the workspace's own documents through <retrieval>, and may
call tools. Answers that use documents cite the source chunks. It must be runnable
end-to-end on day one (even with a mock provider and seeded documents), not a set of screens
wired together later.
Non-goals (do not build unless the must-have list says otherwise): multi-agent orchestration, fine-tuning, a public model marketplace, voice I/O, autonomous background agents.
2. Roles
| Role | Can do |
|---|---|
| End user | Create conversations, chat with streaming answers, upload attachments, see citations and their own usage. |
| Workspace admin | Manage the knowledge base (upload, re-index, delete documents), configure the agent (system prompt, model, temperature, tools) and review usage/cost. |
| Platform operator | Set provider credentials and per-tenant quotas, inspect provider health and the audit log. |
Authentication is required for every conversation and for the whole admin surface; there is no anonymous chat. Provide a seeded admin account and a seeded user for local development.
3. Information architecture (routes)
/— conversations list: newest-first, title, last message snippet, model, new-conversation action./conversations/:conversationId— one conversation: streaming chat, message history, attachments, stop/regenerate, rename/delete./knowledge— knowledge base: document list with status (uploaded → processing → indexed → failed), upload, re-index, delete, per-document chunk count./knowledge/:documentId— one document: metadata, ingestion progress, chunk preview./agent— agent configuration: system prompt, model, temperature, retrieval on/off, tool selection./usage— usage and cost: tokens and cost per period, per conversation and per model./settings— profile, provider keys (BYOK), default model, export/delete account data./login,/logout(admin surface guarded by the admin role).
4. Data model
Minimum viable entities (add fields the requirements imply; keep them typed and validated):
- User — id, email, passwordHash, name, role (
user|admin|operator), createdAt. - Conversation — id, userId, title, model, systemPromptSnapshot, createdAt, updatedAt, archivedAt?.
- Message — id, conversationId, role (
system|user|assistant|tool), content, status (pending|streaming|complete|error|cancelled), inputTokens, outputTokens, model, toolCallId?, createdAt. - Attachment — id, messageId, kind (
image|file), filename, mimeType, sizeBytes, storagePath. - Document — id, workspaceId, filename, mimeType, sizeBytes, status (
uploaded|processing|indexed|failed), error?, chunkCount, uploadedBy, createdAt. - DocumentChunk — id, documentId, ordinal, text, tokenCount, embeddingRef (vector column for
pgvector, else the external collection/point id), createdAt. - AgentConfig — id, workspaceId, systemPrompt, model, temperature, topP, retrievalEnabled, retrievalTopK, enabledToolIds[], updatedAt.
- ToolDefinition — id, workspaceId, name, description, parameters (JSON schema), handler, sideEffecting (boolean), enabled.
- ToolCall — id, messageId, toolId, arguments (JSON), result (JSON), status (
pending|awaiting_confirmation|running|succeeded|failed|denied), latencyMs. - UsageRecord — id, userId, conversationId?, messageId?, model, inputTokens, outputTokens, costMinor, currency, createdAt.
Retrieval storage is chosen in the dialog: pgvector keeps embeddingRef as a vector column;
Qdrant/OpenSearch store the external collection and point id; None (chat only) disables
ingestion and citations. Keep this behind one retrieval interface so the choice is a
configuration, not a rewrite.
Embeddings and model calls always run server-side; the browser never holds a provider key.
5. Key flows
- Streamed reply. The user sends a message; the server persists it, calls the model through
the provider abstraction, and streams tokens back over the response stream. The assistant
message is
streamingwhile tokens arrive andcompletewhen the stream ends. The user can cancel mid-stream (cancelled, partial text kept); a provider error ends the message aserrorwith a retry action. - Document ingestion. The admin uploads a document; the server stores it, extracts text,
chunks it, embeds each chunk, writes it to
<retrieval>, and marks the documentindexed. Progress is visible per document (status + chunk count); a failure records the error and the document stays re-indexable. - Retrieval-augmented answer. When retrieval is enabled, the question is embedded, the top-k
chunks are fetched, and they are passed to the model as untrusted context. The answer includes
citations that map back to the source chunk (document + ordinal), shown next to the
message and linking to
/knowledge/:documentId. - Tool calling. The model may request a tool. Read-only tools run immediately; a
side-effecting tool pauses the turn in
awaiting_confirmationand the user confirms or denies before it executes. The result is fed back to the model and the turn continues; every call and its status is recorded on the message. - Conversation lifecycle. A conversation is auto-titled from its first exchange, can be renamed, and can be deleted (with confirmation) along with its messages and usage links.
6. Functional requirements
- Chat: streaming token-by-token with visible typing state; stop/regenerate; markdown rendering with code blocks; per-message status; retry after an error without losing history.
- Provider abstraction: one interface with
DeepSeek,OpenAI,AnthropicandBYOKimplementations, plus a mock provider used when no key is configured, so the app runs without secrets. - Knowledge base: upload (multiple files), list with status and filters, re-index, delete; extraction for at least text/PDF/markdown; chunk size and overlap documented.
- Retrieval: configurable top-k; citations on every grounded answer; a clear "no relevant documents" state instead of an invented answer; retrieval can be toggled per workspace and per conversation.
- Tools: a registry with JSON-schema parameters; per-workspace enable/disable; the side-effecting confirmation step; timeouts on every handler.
- Usage: record input/output tokens and cost per request; aggregate by period, conversation and model; surface remaining quota to the user.
- Admin: edit the system prompt, model, temperature and tools; guard the surface with the admin role; no destructive action without confirmation.
- Seed data: one workspace, one admin, one user, a few documents already indexed, and a couple of conversations so the app is presentable on first run.
7. Non-functional requirements
- Secrets: never expose model/API keys to the client — all provider calls are server-side and
proxied; keys live in server env only and are documented in
.env.example; BYOK keys are encrypted at rest and never returned to the browser. - Limits: per-user and per-tenant rate limits and token/cost quotas enforced server-side; a request over quota is rejected with a clear message, not a silent failure.
- Untrusted input: documents, tool output and retrieved chunks are untrusted data — never concatenated into instructions, delimited and labelled as data in the prompt, and never allowed to trigger side-effecting tools without the confirmation step.
- Resilience: provider timeouts, rate-limit and partial-stream failures degrade with a clear message and a retry path; a dropped connection keeps the partial answer.
- Context budget: conversation history sent to the model is bounded with a documented truncation/summarisation strategy; the full transcript is always kept server-side.
- Observability: structured logs for model calls, retrieval and tool calls (no prompt or key material in logs); a health endpoint that reports provider reachability.
- Accessibility: semantic landmarks, labelled inputs, keyboard-operable chat, visible focus, contrast at least AA, and an announced streamed message for screen readers.
8. Acceptance criteria (definition of done)
- Install, dev server, lint, typecheck, tests and production build all pass.
- A streamed answer works end-to-end: tokens appear progressively, cancellation stops the stream, and a provider error shows a retry without losing history.
- A question over an ingested document returns a grounded answer with citations to the source chunks.
- An unavailable provider degrades with a clear, actionable message instead of a blank screen.
- Model/API keys are never present in the client bundle or any client response.
- Usage (input/output tokens and cost) is recorded per request and visible in
/usage. - A side-effecting tool cannot run without an explicit user confirmation.
- Seeded data makes the app presentable immediately;
.env.exampledocuments every secret. - README quickstart (install, run, test, env) is accurate; empty/loading/error states exist.
- CI runs install + lint + typecheck + tests + build.
9. Suggested build order
Follow this order and finish (and verify) a layer before starting the next:
- Scaffold the chosen stack, install dependencies, get the dev server and the empty shell running, set up env handling for model keys, and commit the skeleton.
- Streaming chat shell: the message list, the composer and token streaming end-to-end with the mock provider.
- Conversation persistence: conversations and messages in the database, auto-titling, rename/delete, history loading.
- Document ingestion + retrieval: upload → chunk → embed → index over
<retrieval>, and retrieval-augmented answers with citations. - Tool calling: the tool registry, the run loop, and the confirmation step for side-effecting tools.
- Usage + limits: token/cost accounting, quotas and rate limits.
- Quality: tests for the flows above, CI, README,
.env.example, accessibility pass.