Data architecture and engineering, end to end, with AI on top of data you can trust.
Designed, built, governed and handed over
I work across the whole data lifecycle: deciding what the platform should look like, building it, keeping it governed, and making it useful to analysts and models. Most of my work is data architecture and engineering. AI engineering is my second discipline, and it always sits on top of good data.
Target-state designs for lakehouses and warehouses that fit the sources, the team and the budget you actually have, written down as decisions you can defend later.
Typical work
Platform and enterprise data architecture: medallion lakehouse, warehouse and serving layers
Source integration patterns: log-based CDC, high-watermark extracts and dump restores
Technology choices with the tradeoffs on record, such as when a lighter CDC tool serves better than a stream processor
Data models: dimensional marts, slowly changing dimensions and one-table reporting views
Designs that hold up under retries, replays and partial failure, reasoned from how storage engines, logs, replication and transactions actually behave.
Typical work
Design reviews in the vocabulary of Designing Data-Intensive Applications: replication, sharding, transactions, consistency, batch and streams
Change data capture from database logs, ordered by the source's log position rather than arrival time
Exactly-once results from at-least-once delivery: idempotent writes, natural keys, and watermarks committed with the data
Storage engine and table-format choices for the access pattern: B-tree row stores, merge-tree engines such as ClickHouse, and Iceberg tables
Failure analysis that finds the quiet failures, the ones that produce wrong data without an error
Evidence in public work
about 3 seconds from PostgreSQL commit to a queryable ClickHouse row in the reference pipeline
Governance built into the platform rather than added in meetings: contracts at model boundaries, lineage on every run, and access rules enforced at query time.
Typical work
Data contracts and quality gates that stop a bad model before it reaches a dashboard
Lineage with OpenLineage and Marquez, cataloging and PII tagging with OpenMetadata
Role-based access and column masking with Apache Ranger and Trino
Quality controls and pipeline service levels for the datasets the business depends on
Leaving a team able to run what we built: clear handovers between data science and data engineering, documentation that answers the 2 a.m. question, and teaching that sticks.
Typical work
Handover templates that turn a data-science feature request into an engineering spec
Internal lectures on machine learning for analysts and engineers
Runbooks, lineage and tests left in place so the platform does not depend on one person
Community workshops on machine learning, web and mobile development
Evidence for this practice lives on the About page, under community and teaching.
Second discipline
AI and ML work always sits on top of the data practice, never instead of it.
AI and ML on your data
Putting models on top of a platform that can feed them, from feature stores to retrieval over your own documents.
Typical work
Feature stores shared by training and serving, behind low-latency decision APIs
Retrieval-augmented assistants over internal documents, on-premise when data cannot leave
Ethiopian-language NLP: embeddings, OCR and small-model evaluation
Evidence in public work
23% retrieval improvement from EthioLLM embeddings over a multilingual-e5 baseline