Design and implement deterministic detection models for eComms surveillance: market abuse language, insider information patterns, tipping-off phraseology, information barrier breaches, conduct risk, off-channel evasion, and trade-comms correlation
Develop and fine-tune transformer-based NLP models (BERT, RoBERTa, FinBERT) for context-aware detection beyond simple lexicon matching
Build and maintain a model benchmarking framework with ground truth datasets, precision/recall/F1 measurement, AUC-ROC analysis, and automated weekly benchmark runs
Implement threshold tuning workflows to calibrate alert sensitivity per desk, asset class, and jurisdiction — ensuring compliance with ASIC INFO 283 guidance
Develop false negative detection: red team simulated misconduct, historical replay testing, coverage gap analysis, and cross-model ensemble voting
Prepare communication data for transformer models: tokenisation, sequence formatting with context windowing, label engineering from investigations, and data augmentation
Implement model drift detection using PSI and KL divergence with automated alerts when distributions shift
Build champion-challenger evaluation: shadow-mode deployment with statistical significance testing before production promotion
Deliver model explainability using SHAP/LIME for regulatory audit readiness
Produce monthly model effectiveness scorecards for compliance committee review
Collaborate with the eComms pipeline team to ensure clean, normalised inputs for ML model training and inference
Required Qualifications
5+ years in machine learning engineering, with at least 3 years in NLP/text classification in financial services or compliance
Strong proficiency in Python (scikit-learn, PyTorch, TensorFlow, Hugging Face Transformers)
Hands-on experience fine-tuning pre-trained language models (BERT, RoBERTa, GPT-family) for domain-specific tasks
Experience building model evaluation and benchmarking pipelines with automated metric tracking and drift detection