Enterprise AI production platform
From prototype to AI your organisation can rely on.
Aperture sets how much AI enters your organisation, and on what terms. Open-weight models on your own hardware and external API models sit behind one gateway and one set of policies. Every answer is grounded in documents its user is allowed to see, tested before release and monitored after it.
VARDA ApertureRuns on Orbit

The problem
The pilot worked. Production asks different questions.
A model that impressed in a workshop now has to respect document permissions, keep personal data inside, resist manipulated inputs and answer as well next month as it does today. Without a shared platform, each team solves these problems again, differently, and nobody can say which model answered what.
- 01
Teams call external model APIs directly, with no shared policy or cost control.
- 02
Assistants retrieve documents without checking who is allowed to read them.
- 03
Quality is judged by impression rather than by repeatable tests.
- 04
When an answer is wrong, nobody can trace the model, prompt or source behind it.
- Data
- Training
- Evaluation
- Integration
- Deploymentcanary 10%
- Monitoring
- Improvement
- Groundedness0.940.89
- Citation accuracy0.970.93
- Task success0.910.86
- Policy violations0.000.00
- p95 latency1.8 s2.3 s
Awaiting approval: model risk committee
- PII masked (national ID, IBAN)12
- Prompt injection blocked2
- Unsupported answer withheld5
Capabilities
AI enters on your terms.
Model and LLM gateway
One endpoint for open-weight models on your GPUs and API models from external providers. Routing, quotas, fallbacks and cost accounting are set per team and per use case.
Permission-aware retrieval
Retrieval over enterprise documents that applies the source system’s permissions at query time, so an assistant never quotes a document its user could not open.
Evaluation harness
Offline test sets built from real questions are scored for correctness, groundedness and policy breaches. Online evaluation then compares versions on live traffic before a full release.
Guardrails and policies
Personal data is redacted on the way in and on the way out, and prompt-injection attempts are screened in both user input and retrieved content. Output policies block, mask or escalate a response before it reaches anyone.
Versioning and model registry
Prompts, models, retrieval settings and guardrail rules are versioned together as a single release. Promotion to production needs a recorded approval, and every change leaves an audit trail.
Monitoring and feedback
Quality, drift, latency and cost are tracked for each use case. User ratings and expert corrections flow back into the test sets, so every release is measured against what went wrong before.
How it works
One lifecycle, from data to improvement
Data
Sources are connected through Orbit and indexed together with their permissions. Real questions and expert-reviewed answers become the first test set.
Training and adaptation
We start with prompting and retrieval, and fine-tune an open-weight model only where the evaluation shows it will pay off.
Evaluation
Each candidate is scored against the test set and adversarial cases, and release thresholds are agreed with the business owner.
Integration and deployment
The release goes behind the gateway with its guardrails: first in shadow mode, then to a share of users, then to everyone.
Monitoring and improvement
Live quality, drift and cost are tracked; feedback and failures are added to the test set and start the next iteration.
Architecture
VARDA Aperture · Architecture
01 Sources
- Enterprise documents
- Orbit data products
- Open-weight & API models
02 Processing
- Permission-aware retrieval
- Evaluation harness
- Model registry & approvals
03 VARDA Aperture
- Model & LLM gateway
- Guardrails & policies
04 Outputs
- Assistants & applications
- Quality, drift & cost monitoring
- Audit trail
Use cases
Use cases
An illustrative case: operations staff at a bank ask questions across procedures, circulars and product documents. Each answer cites the passage and document version it came from, and only documents the user may open are searched.
- Answers cite the source passage and document version
- Source-system permissions enforced at query time
- Unanswerable questions routed to a subject-matter expert
An illustrative case: an insurer receives claim files as scans, emails and forms. Aperture classifies each document, extracts the required fields against a schema and sends low-confidence cases to a person.
- Structured output validated against a schema
- Confidence thresholds with human review
- Every extraction linked to its model and prompt version
An illustrative case: agents at a telecoms operator receive draft replies grounded in policy documents and the customer’s case history. The agent edits and approves each draft; nothing is sent automatically.
- Personal data masked before any external model call
- Tone and policy rules checked on every draft
- Agent edits feed the evaluation set
Deployment options
Deployment options
On-premises
Open-weight models are served on your own GPU servers, and the vector index, registry and logs stay in your data centre. Nothing leaves the network unless a policy explicitly allows it.
Private cloud
Runs in your own cloud tenancy or a Turkish cloud region on dedicated GPU nodes. API models are reached over private connectivity where the provider offers it.
Hybrid
Sensitive use cases stay on local models, while low-risk ones may be routed to external API models by policy, with personal data redacted first. One gateway and one audit trail cover both.
Technical specification
VARDA Aperture
- Model serving
- Open-weight models on vLLM or NVIDIA Triton; OpenAI-style API endpoint
- Retrieval
- Hybrid BM25 and vector search with reranking; ACLs enforced at query time
- Evaluation
- Offline test suites, human-calibrated LLM judges, shadow and A/B tests
- Guardrails
- PII detection and redaction in Turkish and English; injection filters; output policies
- Versioning
- Prompt, model, retrieval and policy versions pinned per release
- Registry and approvals
- Staged promotion from test to production with recorded sign-off
- Monitoring
- Quality, drift, p50/p95 latency and token cost per use case
- Audit trail
- Prompt, context, response and approver logged, with configurable retention
Integrations
- vLLM
- NVIDIA Triton
- Hugging Face
- OpenAI API
- Azure OpenAI
- MLflow
- pgvector
- OpenSearch
- Qdrant
- SharePoint
- Confluence
- Kubernetes
- OpenTelemetry
- Keycloak
Compliance & governance
- Designed to support KVKK obligations: personal data is detected and masked before it reaches a model, and transfers to external providers can be blocked by policy.
- Designed to support the BDDK information-systems regulation for banks, with models, indexes and logs kept entirely on-premises.
- Designed to support ISO/IEC 27001-style controls for access, change approval and logging across the model lifecycle.
- Designed to support the GDPR principles of data minimisation and accountability where EU residents’ data is processed.

Frequently asked questions
Frequently asked questions
01Does our data leave the building when we use an external model?
Only if you allow it. Each use case has a policy that decides which models it may call, and sensitive ones can be limited to open-weight models on your own hardware. Where an external API is allowed, personal data is redacted first and every request and response is logged.
02Which models can we run?
Open-weight families such as Llama, Mistral and Qwen on your own GPUs, subject to each model’s licence, and commercial models through their APIs. Applications talk to the gateway rather than to a model, so a model can be replaced after evaluation without rewriting the applications that use it.
03How do you decide a release is good enough to go live?
Before release, each use case gets a test set of real questions with expert-reviewed reference answers, scored for correctness, groundedness and policy breaches against thresholds agreed with the business owner. After release, sampled live traffic is scored the same way, and user feedback is added to the test set.
04How do you defend against prompt injection?
In layers. Instructions are kept apart from untrusted content, user input and retrieved passages are screened by classifiers, tools and actions are restricted to allow-lists, and output policies check every response before it is shown. No single defence is sufficient on its own, so each layer is tested with adversarial cases in the evaluation suite.
05How do we keep AI costs under control?
The gateway meters tokens and GPU time per team and use case, with budgets and quotas that alert or throttle when reached. Routine requests go to smaller models, large models are kept for the tasks that need them, and cost is reported next to quality so the trade-off stays visible.
See VARDA Aperture with your own data.
One controlled route for enterprise AI: model gateway, permission-aware retrieval, evaluation, guardrails, monitoring and an audited model registry.





















