Skip to main content
Mati Data

Services

Data architecture and engineering, end to end, with AI on top of data you can trust.

Designed, built, governed and handed over

I work across the whole data lifecycle: deciding what the platform should look like, building it, keeping it governed, and making it useful to analysts and models. Most of my work is data architecture and engineering. AI engineering is my second discipline, and it always sits on top of good data.

Get in touch

The data practice

Data architecture

Target-state designs for lakehouses and warehouses that fit the sources, the team and the budget you actually have, written down as decisions you can defend later.

Typical work

  • Platform and enterprise data architecture: medallion lakehouse, warehouse and serving layers
  • Source integration patterns: log-based CDC, high-watermark extracts and dump restores
  • Technology choices with the tradeoffs on record, such as when a lighter CDC tool serves better than a stream processor
  • Data models: dimensional marts, slowly changing dimensions and one-table reporting views

My production work in this practice is on the CV.

Data-intensive systems design

Designs that hold up under retries, replays and partial failure, reasoned from how storage engines, logs, replication and transactions actually behave.

Typical work

  • Design reviews in the vocabulary of Designing Data-Intensive Applications: replication, sharding, transactions, consistency, batch and streams
  • Change data capture from database logs, ordered by the source's log position rather than arrival time
  • Exactly-once results from at-least-once delivery: idempotent writes, natural keys, and watermarks committed with the data
  • Storage engine and table-format choices for the access pattern: B-tree row stores, merge-tree engines such as ClickHouse, and Iceberg tables
  • Failure analysis that finds the quiet failures, the ones that produce wrong data without an error

Evidence in public work

  • about 3 seconds from PostgreSQL commit to a queryable ClickHouse row in the reference pipeline

My production work in this practice is on the CV.

Data engineering

Production pipelines that fail loudly instead of quietly: orchestrated, tested, idempotent and cheap to rerun.

Typical work

  • Airflow on Kubernetes with retries, data-quality tasks and alerting
  • Change data capture with exactly-once loads into Iceberg and the warehouse
  • dbt on Trino and Redshift, tested and promoted through CI/CD environments
  • Incident work when a production database is under strain: scheduling, restores and recovery

My production work in this practice is on the CV.

Data governance and quality

Governance built into the platform rather than added in meetings: contracts at model boundaries, lineage on every run, and access rules enforced at query time.

Typical work

  • Data contracts and quality gates that stop a bad model before it reaches a dashboard
  • Lineage with OpenLineage and Marquez, cataloging and PII tagging with OpenMetadata
  • Role-based access and column masking with Apache Ranger and Trino
  • Quality controls and pipeline service levels for the datasets the business depends on

My production work in this practice is on the CV.

Analytics engineering and BI

Gold layers that answer the business question in one table, so analysts and dashboards stop joining raw data at query time.

Typical work

  • Reporting masterviews and semantic marts in dbt
  • Serving from Superset, Redshift and ClickHouse with secured self-service access
  • Materialized-view scheduling that keeps operational reporting fast

My production work in this practice is on the CV.

Enablement and training

Leaving a team able to run what we built: clear handovers between data science and data engineering, documentation that answers the 2 a.m. question, and teaching that sticks.

Typical work

  • Handover templates that turn a data-science feature request into an engineering spec
  • Internal lectures on machine learning for analysts and engineers
  • Runbooks, lineage and tests left in place so the platform does not depend on one person
  • Community workshops on machine learning, web and mobile development

Evidence for this practice lives on the About page, under community and teaching.

Second discipline

AI and ML work always sits on top of the data practice, never instead of it.

AI and ML on your data

Putting models on top of a platform that can feed them, from feature stores to retrieval over your own documents.

Typical work

  • Feature stores shared by training and serving, behind low-latency decision APIs
  • Retrieval-augmented assistants over internal documents, on-premise when data cannot leave
  • Ethiopian-language NLP: embeddings, OCR and small-model evaluation

Evidence in public work

  • 23% retrieval improvement from EthioLLM embeddings over a multilingual-e5 baseline

My production work in this practice is on the CV.

How an engagement runs

  1. 1

    Assess

    Map the sources, the current pipelines and where the data goes wrong today. You get a short written assessment with the risks ranked.

  2. 2

    Design

    Propose a target architecture sized to your team and budget, with every significant choice written down as a decision record.

  3. 3

    Build

    Implement the first slice end to end, from one source to one trusted table and dashboard, then widen it source by source.

  4. 4

    Hand over

    Leave tests, lineage and runbooks in place, and pair with your team until they run the platform without me.

Ways to work together

  • Architecture review

    A focused review of an existing platform or a planned design, ending in written recommendations you can act on.

    Good fit: You have a platform or a plan and want a second opinion before committing to it.

  • Platform build

    Designing and building a lakehouse, warehouse or CDC pipeline together with your team.

    Good fit: You need a production data platform and the people to run it afterwards.

  • Governance and quality setup

    Data contracts, lineage, cataloging and access control added to a platform you already run.

    Good fit: Your data exists, but nobody fully trusts it.

  • Senior or lead role

    A full-time role leading data architecture and engineering, remote with a team in Europe or based in Africa.

    Good fit: You are hiring a data architect or a senior data engineer.

Start a conversation

Open to Data Architect and Senior Data Engineer roles, and to selected consulting engagements in data architecture, governance and engineering.