Data operations backbone

Reliable data, from source to decision.

Orbit carries data from the systems that create it to the people and products that depend on it, and makes every step traceable. Ingestion, transformation, quality checks, lineage and access policy run as one platform — and every other VARDA product runs on it.

VARDA Orbit

VARDA Orbit

The problem

Data problems are usually found by the person reading the report.

Pipelines grow one request at a time: a script here, a scheduled export there. Nobody owns the whole chain, schema changes break things silently, and when a figure is questioned it takes days to find where it came from.

  1. 01

    Nightly batch jobs leave operational teams working with yesterday’s data.

  2. 02

    A renamed column upstream breaks reports downstream, and nobody is alerted.

  3. 03

    Personal data is copied into analytics environments without masking or retention rules.

  4. 04

    No one can show, column by column, where a reported figure came from.

VARDA OrbitLineage · finance domainSample interface · illustrative data09:41:18
Column-level lineagelast run 09:33
  • core_banking.oracle
  • erp.sap_s4
  • crm.api
  • iot.mqtt
  • stg_transactions
  • stg_customers
  • stg_telemetry
  • dim_customer
  • fct_transactions
  • fct_machine_hourly
  • feature_store
  • lumen.metrics
  • api.v1
Data contracts & quality
  • ✓not_null(customer_id)
  • ✓unique(transaction_id)
  • ✓freshness < 15 min8 dk
  • !schema drift · erp.sap_s4.MARA +1 columnreview
  • ✓PII · national ID, IBAN masked
Freshness
8 min
Volume today
18.4M rows
Successful runs
99,6%

Capabilities

The orbit that holds the stars.

  • CDC, batch and streaming ingestion

    Log-based change data capture from operational databases, scheduled batch loads and streams from Kafka or MQTT. Schemas are captured on arrival, and raw data is kept so any load can be replayed.

  • Transformation and orchestration

    Transformations are written in SQL and Python, reviewed in Git and tested in CI. The orchestrator runs them in dependency order, with retries, backfills and a service level for every dataset.

  • Data contracts and quality checks

    Producers and consumers agree each dataset’s schema, keys, freshness and allowed values in a contract. Checks run on every load, and data that fails is held back rather than published.

  • Lineage and observability

    Column-level lineage traces every figure from source to dashboard. Freshness, volume and schema drift are monitored per dataset, and when something breaks, lineage shows exactly which reports and models downstream are affected.

  • Governance and catalogue

    Personal data is tagged at ingestion; masking, row-level policies and retention rules are defined once and enforced in every tool. The catalogue records the owner, definition and sensitivity of every dataset.

  • Serving

    Trusted data is published to the warehouse, to REST and GraphQL APIs and to an online and offline feature store, so reports, applications and models all draw on the same governed source.

How it works

How data moves through Orbit

  1. Connect

    Sources are connected through CDC, batch or streaming connectors, and schemas and raw data are captured on arrival.

  2. Contract

    Producers and consumers agree what each dataset must look like: schema, keys, freshness and allowed values.

  3. Transform and test

    SQL and Python models are reviewed, tested in CI and run in dependency order by the orchestrator.

  4. Govern

    Sensitive columns are tagged, masking and retention rules are applied, and lineage records every step.

  5. Serve and observe

    Data is published only when its checks pass, then watched for freshness, volume and schema drift.

Architecture

VARDA Orbit · Architecture

01 Sources

  • Operational databases
  • Applications & files
  • Event streams & IoT

02 Processing

  • CDC & batch ingestion
  • Streaming ingestion

03 VARDA Orbit

  • Transformation & orchestration
  • Contracts, quality & lineage

04 Outputs

  • Warehouse & lakehouse
  • Data APIs
  • Feature store
  • Other VARDA products

Use cases

Use cases

An illustrative case: a retail chain whose stock and sales figures arrive once a night replaces its export scripts with CDC from the ERP and point-of-sale databases. Store operations now work from today’s data, not yesterday’s.

  • Change capture from transaction logs, with minimal load on source systems
  • Full change history kept for audit and replay
  • Downstream tables refreshed continuously, not overnight

Deployment options

Deployment options

  • On-premises

    Runs on Kubernetes or virtual machines in your data centre, close to the source systems. Suited to banks and regulated industries that must keep data in-country and under their own control.

  • Private cloud

    Deployed in your own cloud tenancy or a Turkish cloud region, with object storage for the lakehouse and private links back to on-premises sources.

  • Hybrid

    Ingestion agents run on-premises next to the sources and forward only permitted, masked datasets to cloud analytics. Lineage and policies remain one system across both.

Technical specification

VARDA Orbit

Ingestion modes
Log-based CDC, scheduled batch, streaming (Kafka, MQTT)
Delivery semantics
At-least-once delivery; idempotent writes give effectively exactly-once results
Transformations
SQL and Python models, versioned in Git and tested in CI
Orchestration
Dependency-aware scheduling, retries, backfills and per-dataset SLAs
Quality checks
Schema, freshness, volume, uniqueness and custom rules, per contract
Lineage
Column-level and OpenLineage-based, from source to dashboard
Governance
PII tagging, dynamic masking, row-level security, retention policies
Serving
Warehouse tables, REST and GraphQL APIs, online and offline feature store

Integrations

  • PostgreSQL
  • Oracle Database
  • Microsoft SQL Server
  • SAP S/4HANA
  • Debezium
  • Apache Kafka
  • MQTT
  • Apache Airflow
  • dbt
  • Apache Spark
  • Apache Iceberg
  • ClickHouse
  • OpenLineage
  • Feast

Compliance & governance

  • Designed to support KVKK obligations: personal data is tagged at ingestion, masked by role and retained or deleted according to recorded rules.
  • Designed to support the BDDK information-systems regulation for banks, with in-country deployment, access logs and traceable data changes.
  • Designed to support ISO/IEC 27001-style controls for access management, change control and logging across every pipeline.
  • Designed to support GDPR requirements where EU residents’ data is processed, including records of processing and erasure requests.

Frequently asked questions

Frequently asked questions

01Do we have to replace our existing warehouse or ETL tools?

No. Orbit connects to what you already run: it can land data in your existing warehouse, adopt your SQL models and take over schedules gradually. Legacy jobs move one pipeline at a time, with old and new outputs compared before each switch.

02What happens when a source system changes its schema?

The change is detected on arrival and compared with the dataset’s contract. Additive changes can flow through automatically; breaking changes stop publication of the affected dataset, alert its owner and show, through lineage, which reports and models downstream would be affected.

03How is personal data handled?

Columns are tagged as personal or sensitive at ingestion, by rules and by data-owner review. Masking, pseudonymisation and row-level policies are applied by role, retention periods are enforced by scheduled deletion or anonymisation, and every access and deletion is logged.

04How can we be sure the figures in our reports are right?

Every published dataset has passed the checks in its contract, and lineage shows the exact source columns and transformations behind each figure. When a check fails, the dataset is held back and flagged, so a report shows the last correct data with a warning rather than fresh but wrong numbers.

05Do we need Orbit to use the other VARDA products?

Every VARDA product is built on Orbit and ships with the parts of it that it needs, so there is nothing extra to install. If you already run a mature data platform, those components sit alongside it, reading from your warehouse and adding contracts, lineage and serving only where the products need them.

See VARDA Orbit with your own data.

The data backbone under every VARDA product: CDC, batch and streaming ingestion, tested transformations, lineage, observability and KVKK-aware governance.