Skip to content
View karol-duda's full-sized avatar

Block or report karol-duda

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
karol-duda/README.md

Karol Duda

Senior Financial Crime & Analytics Specialist — FIU investigations · screening & model validation · explainable AI

3+ years across financial crime at Revolut (FIU), PwC (Financial Crime Unit) and AML RightSource: 2,000+ cross-border investigations spanning 12,000+ accounts, SAR/disclosure packages for national FIUs, high-risk CDD/EDD. MSc in Data Science & Society (Tilburg University). I work at the junction the industry struggles to staff: translating financial-crime risk into data, thresholds and governance that survive an auditor.


Featured work

Independent, fully reproducible validation of the open-source screening stack used in production compliance (OpenSanctions nomenklatura / yente) — run the way a model-risk function would.

  • 19.1M live registry records screened (UK PSC + GLEIF vs the EU consolidated sanctions list): 24.5% of review-tier signals are invisible to exact name matching — while exact matching wins on the industry benchmark. The gap is an evaluation artifact, and quantifying it is the core finding.
  • Capacity-calibrated thresholds (recall per alert budget, not F1), segment error analysis (cross-script matches missed 10–25× more often), feature-level decision traces, model card and validation report.
  • 755k-pair benchmark audit: class balance, label provenance, leakage-safe splits, score-saturation diagnosis.

Can you govern a fraud model whose explanations change between retrains even when performance doesn't?

  • 590,540 real e-commerce transactions, strict out-of-time validation, 30-seed stability study across 435 model pairs.
  • Top-5 SHAP feature agreement: 0.692 despite statistically indistinguishable PR-AUC — raised to 0.911 via class weighting and seed ensembling.
  • Conclusion with governance teeth: explanation stability must be validated and monitored separately from predictive performance.

Toolbox

Python SQL pandas scikit-learn XGBoost SHAP multiprocessing at 19M-record scale Power BI Tableau — applied to transaction monitoring, sanctions screening, fraud typologies, EDD/KYC and model-risk documentation.

Contact

LinkedIn · karol.duda113@gmail.com · Polish/EU citizen, unrestricted EU work authorisation

Popular repositories Loading

  1. karol-duda karol-duda Public

  2. dss_thesis_aml_shap dss_thesis_aml_shap Public

    MSc thesis (Tilburg University): measuring and controlling SHAP explanation instability in high-stakes AML/fraud models — 590k real transactions, 30-seed stability study, out-of-time validation

    Jupyter Notebook

  3. sanctions-screening-validation sanctions-screening-validation Public

    Operational validation of the OpenSanctions screening stack: benchmark audit, capacity-calibrated thresholds, segment error analysis, and a 19M-record live-registry test (UK PSC + GLEIF vs EU sanct…

    Jupyter Notebook