I help teams design, build and govern data platforms they can trust.
Senior Data Engineer at Kifiya Financial Technology in Addis Ababa, running lakehouse, CDC and warehouse platforms for digital lending since 2024. I take data work from architecture and governance to production pipelines and analytics. AI engineering is my second discipline.
How I can helpA pipeline I designed, with public code
Select a stage to see the decision behind it and the evidence. Arrow keys move between stages; Enter opens the case study.
- Data flow
- Control plane
Debezium WAL through pgoutput into Redpanda
Debezium reads the write-ahead log through pgoutput over an explicit publication and writes flattened events to Redpanda, with the operation, the LSN and a deleted flag beside each row. The natural key is the primary key, so a delete still names the business entity.
Evidence
about 3 seconds from PostgreSQL commit to a queryable ClickHouse row
- Debezium
- Redpanda
- pgoutput
Read the case study: wb-cdc-analytics: PostgreSQL to ClickHouse CDC in one command
Debezium reads the write-ahead log through pgoutput over an explicit publication and writes flattened events to Redpanda, with the operation, the LSN and a deleted flag beside each row. The natural key is the primary key, so a delete still names the business entity.
Evidence
about 3 seconds from PostgreSQL commit to a queryable ClickHouse row
- Debezium
- Redpanda
- pgoutput
Read the case study: wb-cdc-analytics: PostgreSQL to ClickHouse CDC in one command
Data architecture and engineering, end to end
I work across the whole data lifecycle: deciding what the platform should look like, building it, keeping it governed, and making it useful to analysts and models. Most of my work is data architecture and engineering. AI engineering is my second discipline, and it always sits on top of good data.
Explore the servicesEvidence you can check
Figures from my public code and research. Production figures from my employer work are on the CV, and the full case studies are shared on request.
- 58
- dbt tests and 59 unit tests green in CI in my public CDC pipeline
- 10
- Prometheus alert rules, each with a runbook and a promtool test
- 306
- Ethiopic characters recognised by my Amharic OCR app
- 4-bit
- costs about 3 points, 3-bit about 20 or more, in my Amharic SLM quantization study
Platform credentials
Career tracks completed alongside the production work: the warehouse and lakehouse platforms most data teams run on.
All credentialsDatabricks
Associate Data Engineer in Databricks
DataCamp career track, 2026
Snowflake
Associate Data Engineer in Snowflake
DataCamp career track, 2026
View the statementSQL and PostgreSQL
Associate Data Engineer in SQL
DataCamp career track, 2026
View the statement
Data only helps when people trust it. Here is how I make that happen.
From the first architecture decision to the dashboard your team opens every morning: designed, built, governed, and handed over so it keeps working.
Data architecture
Target-state designs for lakehouses and warehouses that fit the sources, the team and the budget you actually have, written down as decisions you can defend later.
Data-intensive systems design
Designs that hold up under retries, replays and partial failure, reasoned from how storage engines, logs, replication and transactions actually behave.
Data engineering
Production pipelines that fail loudly instead of quietly: orchestrated, tested, idempotent and cheap to rerun.
Data governance and quality
Governance built into the platform rather than added in meetings: contracts at model boundaries, lineage on every run, and access rules enforced at query time.
Analytics engineering and BI
Gold layers that answer the business question in one table, so analysts and dashboards stop joining raw data at query time.
Enablement and training
Leaving a team able to run what we built: clear handovers between data science and data engineering, documentation that answers the 2 a.m. question, and teaching that sticks.
Second discipline
AI and ML on your data
Putting models on top of a platform that can feed them, from feature stores to retrieval over your own documents.
Industry focus
Fintech data, from the inside
My production career has been inside fintech: digital lending, credit scoring, and the data that risk teams, finance teams and partner institutions depend on every day. I know what a repayment schedule, a days-past-due bucket and a credit decision need from a data platform, and what goes wrong when the platform cannot supply it.
See the fintech page- Can we trust today's portfolio-at-risk number?
- How do we add a new lending partner without building a new pipeline?
- Which features can the credit model use in production, and how fresh are they?
- How do we join on personal data without exposing it to analytics?
- We have monthly batch reports. Do we need real-time?
Selected work, and how it was built
Case studies written as constraints, decisions and outcomes. My public projects are open to everyone; the production case studies from my employer work are shared on request.
See all the case studiesOpen source
wb-cdc-analytics: PostgreSQL to ClickHouse CDC in one command
An independent reference implementation of a CDC analytics pipeline: a public REST API into PostgreSQL, Debezium into Redpanda, ClickHouse with dbt marts, Airflow, Prometheus and Grafana, started with one command.
Research
How small, how quantized? Tokenizer, size and quantization for Amharic SLMs
4-bit costs about 3 points, 3-bit about 20 or more
AI engineering
On-premise white-label RAG assistant for regulated institutions
A fully on-premise, white-label assistant for regulated institutions: nginx with TLS, a FastAPI gateway with JWT and Keycloak SSO, PII redaction before inference, AnythingLLM retrieval and Ollama, in Amharic and English.
AI engineering
Amharic OCR for Ethiopic script with HHD-Ethiopic and MMOCR
A full-stack OCR app for Ethiopic script: hierarchical detection, HHD-Ethiopic CTC recognition with confidence scores, multi-page PDFs and a Next.js canvas, plus an MMOCR research line with SATRN and DBNet++.
AI engineering
Natural language to SQL analytics agent over ClickHouse
A prototype that turns plain-English questions into ClickHouse SQL with a Groq-hosted model and returns tables, charts and a short explanation in a React chat interface; built as a technical assessment prototype.
Private case studies
The production platforms from my employer work: lakehouse design, change data capture, governance, a feature store and more. Shared with approved viewers after a short request.
What colleagues say
I've had the pleasure of working closely with Matiwos, a Data Engineer at Kifiya, and his growth has been exceptional. He learns new technologies with impressive speed, consistently applies them effectively, and brings sharp problem-solving skills to every task. His focus, professionalism, and ability to deliver high-quality results make him stand out. Matiwos is an outstanding data engineer with tremendous potential, and I highly recommend him.
I've had the pleasure of working with Matiwos and want to commend him for his exceptional work on our team. He demonstrated outstanding technical proficiency in managing and optimizing data pipelines, ensuring data integrity, and implementing efficient solutions that significantly improved our data processing capabilities. Matiwos is quick to take action, consistently responding to challenges and tasks with enthusiasm and a solutions-oriented mindset. His ability to communicate effectively coupled with strong problem-solving skills, made him an invaluable asset. I am confident that he will bring the same dedication and expertise to any future role.
Get to know me better
About me
How I got here, how I work, and the people I have worked with.
Read about meThe CV
General, data engineering and AI versions, each ready to print as a PDF.
Open the CVWriting
Notes on retrieval, forecasting, data warehousing and Amharic language models.
Read the writing