Mohaiminul

Md Mohaiminul Islam portrait

Md Mohaiminul Islam

Senior Data Scientist · Workers' Compensation Board (WCB), Edmonton

Applied AI that stands up to review.

70M+ documents · ~5K/day scored · ~91% F1 across 14 classes

I build reviewable AI for regulated document work: PII detection, Databricks RAG, and governed releases.

Scroll for proof

Proof at a glance

Fit for regulated AI delivery.

Current work, measured results, and platform depth before the case studies.

Current focus

Applied AI, NLP, and LLM systems

Strongest signal

Reviewable PII detection across WCB's historical claims corpus

Working pattern

Model work paired with Databricks pipelines, evaluation, and human review

Helpdesk virtual agent

Approved for company-wide rollout to ~3,500 employees

Employee-facing Databricks RAG app; a two-week helpdesk-team trial found ~70% of answers helpful, with remaining gaps traced to connecting more internal sources.

Forms workflow POC

Evaluation shows form returns can drop from ~40% to ~10%

Medical report-generation agents with human review; prototype evaluated on historical data, now under departmental review.

Agentic maintenance

Monitors daily production jobs and stages fixes for human approval

Built on Databricks LLM endpoints (Claude), the agentic harness diagnoses failures and validates fixes on sandboxed data samples; later extended to automated code review.

Claim-duration model

3-class XGBoost model in production, scoring ~1K claims/day

Chosen over NLP/Transformer alternatives after benchmarking accuracy, speed, and cost; monitored weekly with retraining alerts.

Annotation automation

Avoided an estimated $25K in manual labeling cost

AWS Lambda workflow for LLM data generation, validation, and annotation at Blue Guardian Canada.

Data preparation

About 2M records prepared; model-prep time down ~30%

Config-driven PySpark preprocessing and embedding pipelines at Servier Canada.

Review loop

From messy documents to reviewed output.

The flagship operating pattern: narrow AI assistance, human review, data iteration, and measured outcomes.

Messy input

Text-heavy forms and internal workflows

PII, support requests, and medical report-generation agents.

Narrow AI assist

Extraction, redaction, and form preparation

LLM work stays scoped to steps staff can check.

Staff review

Human checks remain in the operating path

Outputs are useful only when staff can inspect them.

Evaluated workflow

Form returns can drop from ~40% to ~10%

Historical-data evaluation of the medical report-generation agents, now under departmental review.

Databricks release gatesPII detectionEmployee-facing RAGForms workflow POC

Experience

Experience and education.

Mohaiminul is a Senior Data Scientist in Edmonton, building reviewable LLM and NLP systems for claim-document workflows at WCB. Earlier roles add transformer team leadership, annotation automation, and data-platform work, backed by graduate research in privacy-aware biomedical machine learning.

2023 - Current

Workers' Compensation Board (WCB)

Senior Data Scientist

PII detection at document-corpus scale, Databricks pipelines, employee-facing RAG, agent-assisted forms, and an agentic maintenance harness for regulated claim workflows. Managed 2 interns, mentored 5 junior data scientists, supported 5 hiring processes, and presented recommendations to the CTO and director-level stakeholders.

2023

Quantolio

Senior Data Scientist

Python refactoring, Vision Transformer analysis, and a Streamlit portfolio-analytics prototype with time-series analysis and reinforcement-learning experiments for portfolio optimization.

2023

Blue Guardian Canada Inc.

Lead Data Scientist, ML/NLP

29-class mental-health text classification, synthetic-data strategy, annotation automation, and intern-team leadership.

2021 - 2022

Servier Canada

Data Scientist Intern, Mitacs

Prepared about 2 million unstructured records, cut model-preparation time by ~30%, and supported generative chemistry model work.

2015 - 2023

University of Manitoba

Research Assistant

Privacy-preserving biomedical ML, published research, and evaluation-led modeling across sensitive datasets.

Earlier work includes a data-science internship at Sightline Innovation and teaching in computer science.

Education

  • PhD Candidate in Computer Science (candidacy completed), University of Manitoba
  • MSc in Computer Science, University of Manitoba
  • BSc in Computer Science and Engineering, University of Chittagong

Capability stack

Applied AI, data platforms, and technical leadership.

The practical stack behind the resume and case studies.

Applied AI systems

  • Reviewable LLM workflows
  • PII and entity detection
  • RAG
  • AI agents
  • LangChain
  • LlamaIndex
  • Databricks Vector Search
  • Hugging Face + vLLM
  • Human review loops

Data platforms

  • Azure Databricks
  • PySpark
  • ETL pipelines
  • SQL
  • Unity Catalog
  • Model monitoring
  • Pandas
  • NumPy
  • GCP Vertex AI notebooks

ML engineering

  • Python
  • PyTorch
  • XGBoost
  • Transformers (BERT)
  • Clustering and unsupervised ML
  • Reinforcement learning
  • LLM-as-judge evaluation
  • TensorFlow
  • MLflow
  • DVC
  • FastAPI
  • Docker

Technical leadership

  • 4 direct-report interns (Blue Guardian)
  • Junior data scientist mentoring
  • Cross-functional coordination
  • Reviewable system design

Recognition

Recognition, publications, and platform proof.

Awards, publication-backed research, and platform evidence from the latest resume.

Recognition

01

Prime Minister Gold Medal

02

Research Completion Award

03

VADA Big Data Challenge Winner

Selected publications

2021

A Maximum Flow-Based Approach to Prioritize Drugs for Drug Repurposing of Chronic Diseases

2020

An integrative deep learning framework for classifying molecular subtypes of breast cancer

12 applied AI/ML publications with 200+ citations, including a first-author paper cited 100+ times.

View all on Google Scholar (opens in new tab)

Platform proof

  • Azure Databricks, MLflow, Unity Catalog registry, release gates, audit logging, and token/cost tracking in shipped applied AI work.
  • GCP Vertex AI notebooks, AWS, Azure DevOps, and production-oriented project tooling.

Contact

If staff can't check it, it doesn't ship.

Focus: reviewable LLM/NLP systems, Databricks pipelines, PII detection, and regulated document workflows in Canada.

Md Mohaiminul Islam · Edmonton, Alberta, CanadaApplied AI and NLPRegulated workflows