David Kimanthi

David Kimanthi

Mentor
Rising Codementor
US$15.00
For every 15 mins
ABOUT ME
Senior Data Engineer
Senior Data Engineer

Data Engineer with 8+ years of experience in software and data engineering. I have hands-on experience building scalable cloud data platforms and analytics systems on AWS, GCP, and Plantir Foundry. Expertise in designing reliable ELT pipelines, data warehouses, and data models using Python, Spark, SQL, dbt, and Airflow to support large-scale analytics and machine learning. Proven track record improving pipeline performance and delivering trusted datasets that enable product analytics, experimentation, and data-driven decision-making.

English
Nairobi (+03:00)
Joined October 2024
EXPERTISE
6 years experience
3 years experience
4 years experience
3 years experience
6 years experience
5 years experience
4 years experience

REVIEWS FROM CLIENTS

David's profile has been carefully vetted and approved as a Codementor. Connect with David now, and leave a review for them once you're done!
EMPLOYMENTS
Senior Data Engineer
Proxify
2025-10-01-Present
  • Architected and shipped a production-grade ELT pipeline ingesting 30K+ order line items daily across 12 Amazon marketplaces using...
  • Architected and shipped a production-grade ELT pipeline ingesting 30K+ order line items daily across 12 Amazon marketplaces using SP-API, AWS S3, Snowflake,dbt Core, powering executive revenue dashboards and Finance reporting.
  • Designed a config-driven Apache Airflow framework with dynamic task mapping, region-scoped API rate limiting supporting both incremental and historical data backfill.
  • Developed a dbt Core star schema (staging, intermediate, marts) implementing date-based revenue recognition and dimensional models for sales, refunds, and settlement reconciliation.
  • Partnered with Finance and business stakeholders to define reporting requirements, validate revenue business rules, and iteratively deliver trusted analytics solutions supporting executive decision-making.
Python
SQL
Node.js
View more
Python
SQL
Node.js
Snowflake
Airflow
AWS
Llm building and deployment
Llm agents
Claude code
Agentic system design
View more
Data Engineer III
IDinsight LTD
2022-07-01-2025-07-01
  • Architected and delivered an enterprise-grade data warehouse on AWS RDS (PostgreSQL) supporting 82+ KPIs across 8 national progra...
  • Architected and delivered an enterprise-grade data warehouse on AWS RDS (PostgreSQL) supporting 82+ KPIs across 8 national programs, enabling centralized analytical reporting, scorecards, and data-driven decision-making for cross-functional teams.
  • Eliminated 15+ hours of weekly manual reconciliation of financial data by engineering a daily ELT pipeline (Airbyte, dbt, BigQuery, Looker Studio) that unified data across 3 disparate source systems for a 200+ project portfolio, enabling real-time monitoring and reducing reporting errors for leadership.
  • Built and deployed scalable automated data pipelines using Python (Flask and FastAPI), Apache Airflow, Docker, dbt, and Databricks, enabling near real-time analytics and improving source file processing timeliness and Power BI dashboard refresh performance.
  • Implemented automated testing, CI/CD pipelines with GitHub Actions and Terraform, and version control standards, reducing pipeline failures and improving deployment reliability across data environments.
  • Built web applications to consume API data assets using Node JS, TypeScript, and Fast API
Python
SQL
Node.js
View more
Python
SQL
Node.js
React
Software Development
Apache Spark
Terraform
Data modeling
Data warehouse
Data Pipelines
DBT
AWS
View more
Data Engineer
Hence Technologies
2021-08-01-2022-07-01
  • Built an end-to-end data pipeline in Palantir Foundry using Python, BeautifulSoup, and PySpark to ingest and transform web-scrape...
  • Built an end-to-end data pipeline in Palantir Foundry using Python, BeautifulSoup, and PySpark to ingest and transform web-scraped data, creating a unified master view of legal professional records for a curated marketplace.
  • Designed Foundry ontologies modeling legal entities and applied batch matching methodologies to deduplicate and enrich records, enabling structured search and filtering across thousands of marketplace entries.
  • Developed PySpark transformations to clean, deduplicate, and enrich source data using fuzzy matching and rule-based logic, improving dataset consistency, data fitness for use, and search accuracy across the platform.
  • Implemented ontology write-back workflows to propagate user actions and maintain consistency between the data layer and underlying pipelines, enforcing data quality rules throughout the system.
Python
SQL
NLP
View more
Python
SQL
NLP
JavaScript
Apache Spark
View more
PROJECTS
Apache Spark and Web Scraping
2021
Used Python, Beautiful soup, and Django to web scrape data from the web Used Apache Spark in Palantir Foundry to clean, transform, and lo...
Used Python, Beautiful soup, and Django to web scrape data from the web Used Apache Spark in Palantir Foundry to clean, transform, and load the scraped data into data objects Designed data ontologies to represent real-world business components
Python
Django
Apache Spark
View more
Python
Django
Apache Spark
View more
A Data Warehouse Project on AWS
2023
Created a data warehouse on AWS RDS using dimensional modeling and receiving operational data from several sources Used Airbyte to perfor...
Created a data warehouse on AWS RDS using dimensional modeling and receiving operational data from several sources Used Airbyte to perform data extraction and loading while using Data Build Tool(DBT) to apply SQL transformation for various marts Used Terraform scripts for Infrastructure as Code(IaC) for all the AWS infrastructure The final cleaned data was loaded into data marts that supported PowerBI dashboards
PostgreSQL
ETL
Data modeling
View more
PostgreSQL
ETL
Data modeling
Data warehouse
Apache Airflow
DBT
View more