New Score: a PySpark and neural network credit decisioning engine for Brazil's largest credit bureau
The Problem
Serasa Experian needed a new credit decisioning engine. The existing solution could not keep up with the growing data volume. Processing delays slowed down the credit decisions that depended on it.
Constraints
- Large data volumes with skewed distributions
- Credit decisions for financial institutions across Brazil depended on the output
- A team of 6 engineers to lead through delivery
What I Did
- Led 6 engineers and owned technical direction, sprint planning and code quality
- Designed ETL pipelines on PySpark and Hadoop that process transaction data across distributed clusters
- Optimized join operations and added custom partitioning to handle skewed data
- Delivered New Score, a neural network credit decisioning engine fed by those pipelines
Result
1000s
Companies Using It
6
Engineers Led
New Score is used by thousands of Brazilian companies.
Stack
Python
PySpark
Apache Spark
Hadoop
scikit-learn
PostgreSQL
Data Pipeline That Cannot Keep Up?
Tell me the volume and where it slows down. See how I work on data pipelines.