New Score: a PySpark and neural network credit decisioning engine for Brazil's largest credit bureau

Serasa Experian | Credit bureau, Brazil | Senior Backend Engineer, Tech Lead at Dextra | Nov 2019 to Jul 2020

The Problem

Serasa Experian needed a new credit decisioning engine. The existing solution could not keep up with the growing data volume. Processing delays slowed down the credit decisions that depended on it.

Constraints

  • Large data volumes with skewed distributions
  • Credit decisions for financial institutions across Brazil depended on the output
  • A team of 6 engineers to lead through delivery

What I Did

  • Led 6 engineers and owned technical direction, sprint planning and code quality
  • Designed ETL pipelines on PySpark and Hadoop that process transaction data across distributed clusters
  • Optimized join operations and added custom partitioning to handle skewed data
  • Delivered New Score, a neural network credit decisioning engine fed by those pipelines

Result

1000s
Companies Using It
6
Engineers Led

New Score is used by thousands of Brazilian companies.

Stack

Python PySpark Apache Spark Hadoop scikit-learn PostgreSQL

Data Pipeline That Cannot Keep Up?

Tell me the volume and where it slows down. See how I work on data pipelines.