Skip to content
Dhruv Shah.
Open to software engineering opportunitiesToronto, Canada

Dhruv Shah

Software engineer building reliable systems for genomic and scientific data

I build reproducible data pipelines, machine-learning systems, and automation tools for scientific and operational data, with a particular interest in bioinformatics, genomics, and dependable software infrastructure.

FocusBioinformatics and genomics

EngineeringReliable pipelines and automation

MethodsScientific computing and machine learning

01 / Selected work

Featured projects

Selected work combining biological data, reproducible analysis, machine learning, and dependable data engineering.

Single-cell bioinformatics

01

Stem Cell Differentiation Pipeline

An end-to-end single-cell RNA-seq workflow that processed 50 GB of raw sequencing data from approximately 120,000 cells in under two hours.

Implemented quality-control and normalization steps that recovered the expected transition from stem cells toward heart-muscle cells across day 0 to day 16 samples.

  • R
  • Seurat
  • Nextflow
  • Docker
  • Cell Ranger
  • Python
View repository

Genomics and machine learning

02

Gene Sequence Classifier

A Transformer-based classifier for human, bacterial, viral, and fungal genomic sequences, trained on a custom Biopython encoding pipeline.

Reached 75% classification accuracy and paired the model with a SQL schema for efficient sequence storage and retrieval.

  • Python
  • PyTorch
  • Biopython
  • NumPy
  • SQL
View repository

Reproducible data analysis

03

F1 Sponsor Stock Returns

A Python event-study pipeline that joined Formula 1 race results with market data across 778 sponsor-race observations.

Estimated sub-0.5% next-day movement and near-zero three-day reaction after controlling for sponsor and race fixed effects.

  • Python
  • pandas
  • statsmodels
  • yfinance
  • Matplotlib
View repository

02 / Venture

Building toward personalized cancer care

CancerSync is an early-stage venture exploring how genomic data can become clearer, more actionable, and more useful at the point of care.

Early-stage venture

CancerSync

A platform concept for turning tumor genomics into actionable clinical insight.

View CancerSync

Mission

We are at the early stages of building a platform that will provide oncologists with actionable insights, enabling personalized treatment plans based on the unique genetic profile of each patient’s tumor. Our mission is to improve patient outcomes and push the boundaries of personalized medicine.

What we are working toward

  • Translate complex tumor-genomic data into clear, actionable insights.
  • Support more personalized and evidence-informed treatment planning.
  • Build reliable software at the intersection of oncology, genomics, and clinical decision support.

03 / Experience

Engineering with measurable impact

Production software and data automation work across energy, enterprise systems, and financial operations.

Sep 2025 — Feb 2026

Independent Electricity System Operator (IESO)

Data Analyst — Automation

Built data-analysis and workflow-automation systems for operational and contractual electricity-market processes.

  • Analyzed more than 10 years of data across six hydroelectric generators and flagged payment exceptions affecting $100K.
  • Built documented automation tools that saved eight hours per month and surfaced upcoming deliverables across configurable timeframes.

Jan–Apr 2024; Jan–Apr 2025

KORE Solutions

Software Engineer

Improved the performance, correctness, and test reliability of production enterprise software and data workflows.

  • Reduced an algorithm's runtime by 97%, from five minutes to ten seconds, using a write-optimized hash map.
  • Restored accuracy across 2,000+ records and reduced end-to-end test failures by 80% while cutting CI/CD runtime by 40 minutes.

Jan 2023 — Apr 2023

HTS

Accounting Analyst — Automation & Data

Automated finance operations and reporting with Python, JavaScript, Selenium, and data-visualization tools.

  • Cut invoice-processing time by 50% and manual-entry errors by 70% through browser automation.
  • Reduced ageing-report preparation from two hours to five minutes with Python, pandas, and Matplotlib.

04 / About

Software for complex, consequential data

I am interested in the engineering layer behind scientific discovery: the pipelines, interfaces, validation rules, and data systems that determine whether an analysis is trustworthy, reproducible, and useful.

My combined computer science and business education helps me connect implementation details with research goals, operational constraints, and measurable outcomes.

Education

University of Waterloo

Bachelor of Computer Science

Expected Sep 2027

Computer science, algorithms, software systems, data, and scientific programming.

Wilfrid Laurier University

Bachelor of Business Administration

Expected Sep 2027

Business analysis, finance, strategy, operations, and product thinking.

05 / Technical skills

Tools selected for the problem

A practical stack for scientific analysis, bioinformatics workflows, automation, and reliable software delivery.

Languages

  • Python
  • R
  • Bash
  • SQL
  • C++
  • C
  • JavaScript
  • Java
  • TypeScript

Bioinformatics

  • Seurat
  • Cell Ranger
  • Nextflow
  • Biopython
  • Single-cell RNA-seq

Data and machine learning

  • PyTorch
  • pandas
  • NumPy
  • statsmodels
  • Matplotlib

Engineering tools

  • Git
  • Docker
  • Cypress
  • Next.js
  • Excel
  • Power Query