I get handed the pipeline nobody wants to touch.
Trusted with architecture decisions and cross-team delivery. I take data from raw source to analytics-ready, batch and streaming, modelled properly, with CI/CD wired in so a change ships the same day someone asks for it.
- 124M+records / day
- 75%infra cost cut
- 99.8%pipeline reliability
- 35Ktxn / minute streaming
01 About
Data Engineer with 5+ years in banking and property tech, owning architecture decisions and leading delivery across teams. In practice the work is the jobs that fail silently, the lineage nobody can draw anymore, the bill that quietly tripled.
In property tech that meant a valuation pipeline that was quietly dropping 19 out of every 20 rows. In banking it meant 17+ sources folded into one datamart, a two-hour lead ETL cut to five minutes, and fuzzy matching over 2M+ records at 30% better accuracy.
Most pipelines break quietly. I build the ones that don’t, and most of my work is measured in what it removed: cost, latency, manual steps, and the phone call at 3 a.m.
02 Experience
-
Sep 2025 — Present
Data Engineer — Big Data, Streaming & Analytics
Bank Negara Indonesia
- Shipped a production Kafka + Spark Streaming pipeline handling 25–35K transactions/minute, with schema registry enforcing schema evolution and data contracts.
- Delivered a unified datamart for 20+ business and campaign reports, cutting query turnaround from a full day to under 5 minutes over 134M+ daily records.
- Built HC data architecture consolidating 17+ sources into a 5-layer medallion datamart, 10+ entity tables for AI-ready analytics on ~36K employees.
- Held production Spark workflows at 99.8% reliability while improving job performance up to 92%.
Kafka · Spark Streaming · PySpark · Hive/Impala · Schema Registry
-
May 2023 — Nov 2025
Data Engineer — Lead Automation & Campaign Analytics
Bank Negara Indonesia
- Automated the sales-lead ETL end to end (passworded Excel, FTP delivery, email alerts), taking processing from 2 hours to under 5 minutes.
- Optimized fuzzy matching on 2M+ records with RapidFuzz + multiprocessing: accuracy +30%, runtime −50%.
- Built a self-healing SQL retry layer — failure resolution time −80%, 100% query success without manual watching.
- Owned delivery and mentored junior and vendor engineers to shorten onboarding.
Python · PySpark · RapidFuzz · SQL · Bash
-
May 2021 — Apr 2023
Data Engineer — Data Marts & Query Optimization
Bank Negara Indonesia
- Automated daily/weekly/monthly datamart refreshes via CDSW — 100% on-time delivery, zero manual runs.
- Partitioned Hive/Impala tables on Parquet: 3x faster queries on less storage.
- Tuned slow SQL across Hadoop, cutting execution time up to 70% and serving 10+ business domains.
Hive · Impala · Parquet · PySpark · Cloudera
Before data: civil engineering & construction — quantity surveyor, structural/finishing inspector, drafter (2016–2020) · STEM teacher, part-time (2013–2014).
03 Projects
Orchestration Migration: Composer → Windmill
Consolidated scattered Airflow DAGs into one Windmill execution model, cutting orchestration infra cost 75% (US$1,200 → US$300/mo).
Windmill · dbt · BigQuery · Python · GCP
Property Valuation Pipeline Rescue
Diagnosed and rebuilt a broken home-valuation ingest — completeness from ~2 to 38 rows/day (~19x), lifting downstream valuation accuracy.
Python · curl_cffi (anti-bot) · BigQuery · Windmill
Medallion Warehouse on BigQuery
Re-architected fragmented sources into a dbt-modeled 5-dataset medallion warehouse; worst-case root-cause debugging from 8 hours to under 15 minutes.
dbt · BigQuery · Medallion architecture
Lead-Matching Lineage Rebuild
Untangled a MATCH lead-matching lineage in 4 days: 113 → 69 stages, downstream consumers from a dozen+ down to 4.
dbt · BigQuery · SQL
Banking Loan Risk ELT
How do banks screen millions of loan records daily? End-to-end pipeline from raw files to risk-ready tables.
Airflow · PySpark · dbt · BigQuery · Docker
E-Commerce ELT on Cloud Composer
Production-shaped Airflow 2.x pipeline with CI/CD and Slack alerting, runnable end to end.
Cloud Composer · dbt · BigQuery · GitHub Actions
Three-Source Sales ELT
Extracts from API, Postgres and Cloud Storage into one modeled sales summary.
Airflow · Docker · dbt · GCP
Banking Streaming OLTP → OLAP
Real-time transaction stream landed into analytics storage without breaking schema contracts.
Kafka · Spark Streaming · Schema Registry
Massive Lead Assignment
Distributes 100K–1M leads across sales reps by geography and customer criteria, fairly, round-robin.
Pandas · PySpark · Hive · CML
High-Speed Fuzzy Matching
Millions of messy name records matched in parallel instead of overnight.
RapidFuzz · Multiprocessing · Python
04 Skills
Languages
Python · SQL · Bash
Orchestration
Airflow · Windmill · dbt · Dagster · Cloud Composer · Cron
Streaming & Big Data
Kafka · Spark Streaming · PySpark · Hadoop · Hive · Impala · HDFS · Databricks
Cloud & Warehouse
BigQuery · Cloud Storage · Cloud Run · GKE · Vertex AI · Snowflake · Medallion architecture · Parquet · Partitioning
DevOps & CI/CD
Docker · Kubernetes · GitHub Actions · Bitbucket Pipelines · Linux · Git
Databases
PostgreSQL · MySQL · Oracle Exadata · Teradata · HiveQL · ImpalaQL
Scraping & Automation
curl_cffi (anti-bot) · Scrapy · BeautifulSoup · Playwright · Selenium · RapidFuzz
Education — Data Science & Machine Learning, Purwadhika Digital Technology School (2020) · B.Eng. Civil Engineering, Pembangunan Jaya University (2013–2017)
05 Contact
If you’ve got a pipeline nobody wants to touch, I’d like to hear about it. Open to full-time or contract work, remote, hybrid, or on-site.