Raw in. Value out.
Aleksa Milic
I build and run data platforms end to end: Spark and Azure Data Factory pipelines, warehouses that stay fast at billions of rows, and the CI/CD and infrastructure under them.
02 / Selected work
Things I’ve built
- 022026
When do you actually need a cluster?
A reproducible TPC-H benchmark: Apache Spark (local, plus a real 2-worker cluster) against DuckDB, Polars and DataFusion on identical Parquet. How far does one machine get before a cluster earns its overhead?
- Role
- Engine adapters, the timing and memory harness, a cross-run correctness gate, the Dockerised cluster, CI, and the report generator.
- Outcome
- The in-process engines finish ~10x faster than Spark at SF1 and SF10, and the gap holds as data grows. Every run gated on 96 correctness checks.
- Apache Spark
- DuckDB
- Polars
- DataFusion
- Docker
- TPC-H
03 / Skills & experience
What I work with
I’m a data engineer with four years building pipelines and cloud data platforms across AWS, Azure, and GCP, mostly on Databricks and Azure Data Factory. Enough of my time has gone into DevOps (containers, CI/CD, infrastructure-as-code) that I can take a pipeline from ingestion to production without handing it off.
Ingestion & orchestration
- Azure Data Factory
- Apache Airflow
- AWS Glue
- AWS DMS
- Python
Processing
- Apache Spark
- PySpark
- Databricks
- Microsoft Fabric
Storage & modelling
- Delta Lake
- Medallion architecture
- AWS Redshift
- Azure Synapse
- BigQuery
- Advanced SQL
Platform, DevOps & BI
- Docker
- Kubernetes
- Terraform
- Multi-cloud CI/CD
- Power BI
- Jan 2026 – Present
Data Engineer
Ingsoftware / ASML
- Run distributed processing on Databricks and Spark over 10+ TB, with Delta Lake for ACID transactions on AWS. Extended the same patterns to GCP (BigQuery, Dataflow) for a multi-cloud setup.
- Cut average runtime on critical Spark jobs by 40% through partition tuning, caching, and cluster right-sizing.
- Own data governance across Databricks and Delta Lake: quality checks, access control, and lineage for production datasets.
- Jun 2024 – Jan 2026
Data Engineer
Vega IT
- Built and ran 20+ ETL/ELT pipelines on Azure Data Factory and Airflow, with reconciliation and data-quality controls gating every load.
- Architected warehouses on AWS Redshift and Databricks Delta Lake over 5+ billion records: partitioning, distribution keys, and incremental models sized to the query patterns.
- Shipped 25+ Power BI dashboards and ran the platform DevOps: ECS/EKS with Docker and Kubernetes, Terraform IaC, multi-cloud CI/CD.
- Jul 2022 – Jun 2024
Data Analyst
Gemini Software
- Automated transformation and ingestion workflows in Python (Pandas, NumPy), replacing steps that had been run by hand.
- Handled cleansing and collation from internal and external sources. Wrote complex SQL with advanced joins, window functions, and CTEs.
- Integrated third-party APIs and contributed to AWS/Azure cloud-migration work.
Education
- 2023 – 2025
MSc, Data Science and Engineering
Faculty of Electronic Engineering, Niš
GPA 9.7 / 10. Thesis: optimising distributed data-pipeline architectures on cloud platforms.
- 2019 – 2023
BSc, Computer Science and Informatics
Faculty of Electronic Engineering, Niš
GPA 8.7 / 10. Thesis: benchmarking distributed data-processing frameworks for large-scale analytics.
04 / Contact
Get in touch
I’m at Ingsoftware these days, on ASML’s data platform. Open to work on data platforms, Spark performance, or multi-cloud. Email is the quickest way to reach me.
aleksa@aleksa-milic.comDownload CVPDF© 2026 Aleksa Milic