Skip to content
View keroloshany47's full-sized avatar

Highlights

  • Pro

Block or report keroloshany47

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
keroloshany47/README.md

Hi there, I'm Kerolos hani

Data Engineer | Passionate about Big Data


About Me

I am a Data Engineer with a strong focus on Big Data systems and scalable data pipelines, alongside a solid background in data analysis.

I specialize in building end-to-end data solutions including data ingestion, processing, and orchestration using tools such as Python, SQL, Apache Spark, Kafka, Airflow, and Databricks. I also work on transforming complex datasets into meaningful insights and dashboards using Power BI.

While I have experience in data analysis, my primary interest lies in designing and optimizing large-scale data engineering systems that enable reliable and efficient data-driven decision-making.


Tech Stack & Tools

Programming Languages

C C++ Python SQL

Data Analysis & Visualization

Power BI Pandas Matplotlib

Data Engineering & Big Data Tools

Apache Spark Apache Kafka Apache Hadoop Apache Airflow Databricks Microsoft Fabric dbt

Tools & DevOps

Git GitHub Docker Linux


Connect with Me


Always open to collaboration, learning, and building impactful data-driven systems

Pinned Loading

  1. StreamFlow StreamFlow Public

    A fully containerized real-time financial transaction streaming pipeline — built to simulate, process, detect anomalies, store, and monitor high-throughput event data at scale.

    Java 2

  2. FraudLens FraudLens Public

    streaming pipeline that detects financial fraud in under 5 seconds, processing 1.8M transactions across real-time and historical paths using Kafka, Spark, Airflow, dbt, and Grafana.

    Python

  3. SQL_Problems SQL_Problems Public

    This repository documented my journy solving SQl problems in multi websites like LeetCode and Hacker Rank

    SQL

  4. Informatica_ETL Informatica_ETL Public

    This repo contain presentation of an ETL using informatica power center

  5. snowflake_dbt_airflow_project snowflake_dbt_airflow_project Public

    This project demonstrates a modern data engineering pipeline ingests raw CSV files, loads them into Snowflake, transforms the data using dbt models, and schedules the execution using Airflow DAGs.

    TSQL 1

  6. AWS_Youtube_Pipeline AWS_Youtube_Pipeline Public

    This project is a serverless AWS pipeline that pulls YouTube trending video data for 10 regions, cleans and validates it, then aggregates it into analytics-ready tables — all coordinated by a singl…

    Python