Enterprise AI production platform

From prototype to AI your organisation can rely on.

Aperture sets how much AI enters your organisation, and on what terms. Open-weight models on your own hardware and external API models sit behind one gateway and one set of policies. Every answer is grounded in documents its user is allowed to see, tested before release and monitored after it.

VARDA ApertureRuns on Orbit

VARDA Aperture

The problem

The pilot worked. Production asks different questions.

A model that impressed in a workshop now has to respect document permissions, keep personal data inside, resist manipulated inputs and answer as well next month as it does today. Without a shared platform, each team solves these problems again, differently, and nobody can say which model answered what.

  1. 01

    Teams call external model APIs directly, with no shared policy or cost control.

  2. 02

    Assistants retrieve documents without checking who is allowed to read them.

  3. 03

    Quality is judged by impression rather than by repeatable tests.

  4. 04

    When an answer is wrong, nobody can trace the model, prompt or source behind it.

VARDA Aperturecontract-assistant · v3Sample interface · illustrative data09:41:18
  1. Data
  2. Training
  3. Evaluation
  4. Integration
  5. Deploymentcanary 10%
  6. Monitoring
  7. Improvement
Evaluation · v3 (candidate) / v2 (production)
  • Groundedness0.940.89
  • Citation accuracy0.970.93
  • Task success0.910.86
  • Policy violations0.000.00
  • p95 latency1.8 s2.3 s

Awaiting approval: model risk committee

Drift · last 30 daysPSI 0,08
  • PII masked (national ID, IBAN)12
  • Prompt injection blocked2
  • Unsupported answer withheld5

Capabilities

AI enters on your terms.

  • Model and LLM gateway

    One endpoint for open-weight models on your GPUs and API models from external providers. Routing, quotas, fallbacks and cost accounting are set per team and per use case.

  • Permission-aware retrieval

    Retrieval over enterprise documents that applies the source system’s permissions at query time, so an assistant never quotes a document its user could not open.

  • Evaluation harness

    Offline test sets built from real questions are scored for correctness, groundedness and policy breaches. Online evaluation then compares versions on live traffic before a full release.

  • Guardrails and policies

    Personal data is redacted on the way in and on the way out, and prompt-injection attempts are screened in both user input and retrieved content. Output policies block, mask or escalate a response before it reaches anyone.

  • Versioning and model registry

    Prompts, models, retrieval settings and guardrail rules are versioned together as a single release. Promotion to production needs a recorded approval, and every change leaves an audit trail.

  • Monitoring and feedback

    Quality, drift, latency and cost are tracked for each use case. User ratings and expert corrections flow back into the test sets, so every release is measured against what went wrong before.

How it works

One lifecycle, from data to improvement

  1. Data

    Sources are connected through Orbit and indexed together with their permissions. Real questions and expert-reviewed answers become the first test set.

  2. Training and adaptation

    We start with prompting and retrieval, and fine-tune an open-weight model only where the evaluation shows it will pay off.

  3. Evaluation

    Each candidate is scored against the test set and adversarial cases, and release thresholds are agreed with the business owner.

  4. Integration and deployment

    The release goes behind the gateway with its guardrails: first in shadow mode, then to a share of users, then to everyone.

  5. Monitoring and improvement

    Live quality, drift and cost are tracked; feedback and failures are added to the test set and start the next iteration.

Architecture

VARDA Aperture · Architecture

01 Sources

  • Enterprise documents
  • Orbit data products
  • Open-weight & API models

02 Processing

  • Permission-aware retrieval
  • Evaluation harness
  • Model registry & approvals

03 VARDA Aperture

  • Model & LLM gateway
  • Guardrails & policies

04 Outputs

  • Assistants & applications
  • Quality, drift & cost monitoring
  • Audit trail

Use cases

Use cases

An illustrative case: operations staff at a bank ask questions across procedures, circulars and product documents. Each answer cites the passage and document version it came from, and only documents the user may open are searched.

  • Answers cite the source passage and document version
  • Source-system permissions enforced at query time
  • Unanswerable questions routed to a subject-matter expert

Deployment options

Deployment options

  • On-premises

    Open-weight models are served on your own GPU servers, and the vector index, registry and logs stay in your data centre. Nothing leaves the network unless a policy explicitly allows it.

  • Private cloud

    Runs in your own cloud tenancy or a Turkish cloud region on dedicated GPU nodes. API models are reached over private connectivity where the provider offers it.

  • Hybrid

    Sensitive use cases stay on local models, while low-risk ones may be routed to external API models by policy, with personal data redacted first. One gateway and one audit trail cover both.

Technical specification

VARDA Aperture

Model serving
Open-weight models on vLLM or NVIDIA Triton; OpenAI-style API endpoint
Retrieval
Hybrid BM25 and vector search with reranking; ACLs enforced at query time
Evaluation
Offline test suites, human-calibrated LLM judges, shadow and A/B tests
Guardrails
PII detection and redaction in Turkish and English; injection filters; output policies
Versioning
Prompt, model, retrieval and policy versions pinned per release
Registry and approvals
Staged promotion from test to production with recorded sign-off
Monitoring
Quality, drift, p50/p95 latency and token cost per use case
Audit trail
Prompt, context, response and approver logged, with configurable retention

Integrations

  • vLLM
  • NVIDIA Triton
  • Hugging Face
  • OpenAI API
  • Azure OpenAI
  • MLflow
  • pgvector
  • OpenSearch
  • Qdrant
  • SharePoint
  • Confluence
  • Kubernetes
  • OpenTelemetry
  • Keycloak

Compliance & governance

  • Designed to support KVKK obligations: personal data is detected and masked before it reaches a model, and transfers to external providers can be blocked by policy.
  • Designed to support the BDDK information-systems regulation for banks, with models, indexes and logs kept entirely on-premises.
  • Designed to support ISO/IEC 27001-style controls for access, change approval and logging across the model lifecycle.
  • Designed to support the GDPR principles of data minimisation and accountability where EU residents’ data is processed.
A long data-centre aisle between rows of server racks, blue status lights reflected in a polished floor.
The same aisle redrawn as a luminous wireframe over a grid floor, with circuit traces branching out from the racks.

Applied domain · A5

Custom AI Systems

From experiment to production

Explore the solution

Frequently asked questions

Frequently asked questions

01Does our data leave the building when we use an external model?

Only if you allow it. Each use case has a policy that decides which models it may call, and sensitive ones can be limited to open-weight models on your own hardware. Where an external API is allowed, personal data is redacted first and every request and response is logged.

02Which models can we run?

Open-weight families such as Llama, Mistral and Qwen on your own GPUs, subject to each model’s licence, and commercial models through their APIs. Applications talk to the gateway rather than to a model, so a model can be replaced after evaluation without rewriting the applications that use it.

03How do you decide a release is good enough to go live?

Before release, each use case gets a test set of real questions with expert-reviewed reference answers, scored for correctness, groundedness and policy breaches against thresholds agreed with the business owner. After release, sampled live traffic is scored the same way, and user feedback is added to the test set.

04How do you defend against prompt injection?

In layers. Instructions are kept apart from untrusted content, user input and retrieved passages are screened by classifiers, tools and actions are restricted to allow-lists, and output policies check every response before it is shown. No single defence is sufficient on its own, so each layer is tested with adversarial cases in the evaluation suite.

05How do we keep AI costs under control?

The gateway meters tokens and GPU time per team and use case, with budgets and quotas that alert or throttle when reached. Routine requests go to smaller models, large models are kept for the tasks that need them, and cost is reported next to quality so the trade-off stays visible.

See VARDA Aperture with your own data.

One controlled route for enterprise AI: model gateway, permission-aware retrieval, evaluation, guardrails, monitoring and an audited model registry.