Context
Regulated institutions want an assistant over their own documents, and the one thing they cannot do is send customer conversations to someone else's servers. I built a white-label assistant that runs entirely on the institution's own hardware: the model, the retrieval, the identity provider and the audit log. One block of configuration, the brand name, slug and domain, re-tenants the whole stack, so the same code serves any institution without edits. It answers in Amharic and English.
I built it alone on my own laptop GPU. The public version is a working demo with sample documents for a fictional institution.
Constraints
- Nothing leaves the building: no external model API, no hosted vector store, no third-party identity service.
- Compliance by construction rather than by policy: raw question text is never stored, personal data is redacted before the model sees it, and every request leaves an audit row.
- Role awareness: customers, front-line staff, managers and administrators each see only their own document set.
- Bilingual without two deployments.
- Re-brandable by configuration, including the sign-in realm and the certificates.
Architecture
Open WebUI is the chat interface. nginx terminates TLS and applies per-IP rate limits on the chat and sign-in routes. Behind it, a FastAPI gateway validates the JWT, issued either locally or by Keycloak with role-mapped claims, runs the question through PII redaction, maps the caller's role to an AnythingLLM workspace and forwards the scrubbed text. AnythingLLM retrieves over the documents in that workspace and calls Ollama, which runs on the host GPU outside Docker. PostgreSQL stores users and audit rows, Redis caches sessions, and Keycloak provides single sign-on.
Redaction covers phone numbers, account numbers, tax identification numbers (TINs), national ID numbers, birth dates and email addresses. The audit row keeps a one-way hash of the original text, the scrubbed text, the workspace, the role at query time, whether personal data was detected and the response time. The original question is never stored.
Generation uses a multilingual instruct model from the Qwen 2.5 family, with a local embedding model for retrieval; both can be swapped in AnythingLLM. The sign-in realm, container names, workspace slugs and certificate subject are rendered from the brand block by two scripts.
Decisions and tradeoffs
- A gateway in front of the retrieval engine instead of modifying it. Authentication, redaction, role mapping and audit live in one small codebase I own; the cost is an extra hop and configuration in two places.
- Redact before inference and log hashes, not text. The model never sees the personal data the patterns catch, and the audit log cannot leak it. The cost is that redaction is pattern-based in this phase: it will miss unusual formats, and it can mangle a legitimate number inside a question.
- One workspace per role rather than document-level permissions. Simple to reason about and to audit; coarse, because a document is either in a workspace or not.
- Ollama on the host, not in a container, for direct GPU access. One component sits outside Compose and is documented separately.
- A small local model over a hosted frontier model. Sovereignty wins and quality loses, Amharic quality most of all. My research on quantizing Amharic language models is about exactly this trade.
Outcome
A working demo: sign in, ask a question, get an answer grounded in the sample documents, and watch the redaction and audit guardrails fire on a question that contains personal data. The launch post is linked below. The gaps are real, and I would rather state them than let the README imply otherwise: there is no streaming (the request schema has a stream flag, but the gateway returns complete responses), there are no automated tests yet, development runs on self-signed certificates, and redaction is regular expressions only. I publish no usage numbers for it.
What I would change
- Tests before features: redaction unit tests with Amharic and English fixtures, token validation tests, and one integration test through nginx.
- Streaming responses through the gateway.
- Named-entity redaction for names and addresses alongside the patterns.
- Document-level permissions on top of the role workspaces.
- A small Amharic evaluation set, so a model swap is measured rather than felt.
- Metrics for Prometheus next to the existing health endpoint.
Stack and links
Open WebUI, nginx, FastAPI, Keycloak, AnythingLLM, Ollama, PostgreSQL, Redis, Docker Compose and Python. The source is private; the launch post is linked below.