Lea Brody-Heine

Software Engineer

MSc Computer Science (AI/ML Focus)

University of St Andrews MSc
Brown University BA

Contact me
Portfolio
 Lea Brody-Heine

My Tech Stack

Anaconda Angular AWS Azure Bootstrap CSS3 D3.js Django Docker Firebase GitHub GitLab HTML5 Java JavaScript Jira scikit-learn Git Jupyter Keras Linux Matplotlib MongoDB MySQL Next.js NumPy Pandas Python PyTorch React TensorFlow TypeScript Vue.js

ABOUT

lines

I’m a Software Engineer at Google, working on Ads/AdMob EngProd, where I build automation and test infrastructure across Android and iOS. I hold a BA from Brown University and an MSc in Computer Science from the University of St Andrews, graduating with First Class Honours with Distinction. Throughout my studies and career, I’ve sharpened my skills in software development, release engineering, data science, and AI, with hands-on experience across Python, Java, JavaScript, TypeScript, SQL, Go, and a range of full-stack and machine learning technologies.

I’m always eager to learn and innovate, constantly seeking opportunities to expand my skill set and make a meaningful impact in the tech industry and beyond. My Master’s dissertation applied machine learning and computer vision to pathology research in mast cell diseases, and I continue to explore how AI and automation can solve real-world problems. I’m excited to keep applying my expertise and contribute to groundbreaking technologies as my career grows.

Download my CV

What I Do

lines
Full Stack Developer

Full Stack Development

  • Proficient in responsive, interactive front ends using HTML5, CSS3, JavaScript, TypeScript, and frameworks like React, Angular, and Vue.js.
  • Strong back-end experience in Java, Python, and Node.js/Express, with RESTful API design and databases including MySQL, MongoDB, and Firebase.
  • Built and launched a full-stack site end-to-end as a paid contractor, from wireframes to deployment, improving discoverability 21% through SEO best practices.
  • Skilled in building responsive, accessible UIs with Bootstrap and Figma.
  • Utilized Git/GitHub for version control and JIRA for project management.
Machine Learning & AI

ML, AI, & Data Science

  • Expertise in Python with ML libraries including TensorFlow, PyTorch, Keras, scikit-learn, and XGBoost for classification and deep learning.
  • Built an end-to-end computational biology pipeline for my Master's dissertation, applying clustering, hypothesis testing, and classifier tuning to microbiome and biopsy data, achieving 91%+ accuracy on a clinical dataset.
  • Designed generative AI (GAN-based) and traditional data augmentation strategies to work around a severely data-limited clinical dataset.
  • Applied computer vision (YOLOv8) and biomedical image analysis to detect and quantify cells in histology/biopsy images, improving detection accuracy by 25%.
  • Built an agentic, Gemini-powered pipeline at Google that autonomously analyzed system logs to surface failure points across 5 product lines.
Mobile, Test Automation, & Release Engineering

Mobile Dev, Test Automation, & Release Engineering

  • Design and build large-scale test automation frameworks and CI/CD pipelines for mobile platforms, reducing manual QA effort and release risk.
  • Drive release-engineering initiatives end-to-end, from technical design docs through implementation and rollout.
  • Build performance-monitoring and observability tooling (latency, memory, and stability metrics) to catch regressions before they reach users.
  • Automate bug triage and test-result analysis at scale, using scripting to cut manual backlog work and speed up debugging.
  • Comfortable owning cross-platform mobile infrastructure end-to-end, from architecture and physical-device testing to on-call support.

Work Experience

lines
AIM Executive Coaching
Sep – Oct 2024
Full Stack Web Developer
  • Built a fully responsive site end-to-end using HTML, CSS, JavaScript, and HubSpot.
  • Increased discoverability 21% (SEO) by implementing best practices, optimizing load times, and supporting mobile, tablet, and desktop.
  • Partnered closely with the client to align the build with wireframes and evolving feedback.
Skills: HTML CSS JavaScript HubSpot CMS SEO Responsive Design Client Communication
GSI Water Solutions, Inc.
May – Aug 2024
Software Engineer Intern
  • Designed and launched a full-stack solution, managing a 10,000+ line codebase integrated with existing infrastructure for timely notifications.
  • Increased client follow-up efficiency 36% with a calendar/email alert system and automated renewal reminders.
  • Owned the product lifecycle end-to-end, from requirements gathering through deployment, as sole developer.
Skills: Full-Stack Development Systems Integration Automation Product Lifecycle Management Independent Ownership

Master's Dissertation

lines

Machine Learning for Pathology in Mast Cell Diseases

University of St Andrews · Jan 2024 – Aug 2024

Download Dissertation

Overview

Mast cell diseases — including hereditary alpha-tryptasemia (HαT) and idiopathic mast cell activation syndrome (i-MCAS) — are notoriously difficult to diagnose, often relying on symptom tracking and treatment response rather than clear biomarkers.

My MSc dissertation at the University of St Andrews explored whether machine learning could help close that gap, applying ML techniques to both pathological clinical data and microbiome data to uncover diagnostic patterns and potential biomarkers for these conditions.

Electron micrograph of a mast cell this is a mast cell! :)
Figure 1. Mast cell (electron micrograph). Provided by Mariana Castells, MD, PhD. Source: tmsforacure.org

Abstract

In this dissertation, we applied machine learning techniques to improve the diagnosis and research of mast cell diseases, focusing on idiopathic mast cell activation syndrome (i-MCAS) and hereditary alpha-tryptasemia (HαT). We developed and integrated machine learning models that utilized both pathological tabular data and biopsy images to build a comprehensive proof-of-concept tool. To overcome the challenges of limited datasets, we compared the success of traditional data augmentation and generative models, followed by the training and testing of various machine learning algorithms. Despite identifying some limitations in generating accurate synthetic data, the traditional augmentation method proved more successful. To further explore i-MCAS, I created a tool to analyze microbiome data and investigate microbial discrepancies in i-MCAS. Our proof-of-concept tool demonstrates the potential for machine learning to discover novel patterns in complex pathological data, though we acknowledge the limitations in collecting and integrating image and tabular data. Further refinement and integration into existing diagnostic frameworks is necessary for broader adoption.

Approach

Built a standardized preprocessing and EDA pipeline in Python to clean, impute, and engineer features from a small, messy real-world clinical dataset (73 observations, 21 raw features), using domain-informed imputation strategies grounded in published clinical ranges rather than naive defaults.

Compared two data augmentation strategies to address the dataset's small size: a traditional noise-based approach and a GAN-based synthetic data generator. Validated both using Kolmogorov-Smirnov distribution testing, and found that the simpler traditional method preserved real data patterns more reliably than the GAN.

Trained and tuned multiple classifiers (XGBoost, MLP, Random Forest, Gradient Boosting, SVM, Logistic Regression) to distinguish between HαT, MCAS, and normal diagnoses, benchmarking models trained on the original dataset against those trained on the augmented dataset.

Designed and built an end-to-end microbiome analysis pipeline using Human Microbiome Project data as a base: generated synthetic i-MCAS and control cohorts via bootstrapping, then applied K-Means and Gaussian Mixture Model clustering to uncover candidate bacterial biomarkers, cross-validating findings with t-tests.

Built the model application layer (script.py) to standardize preprocessing and apply trained models to new patient data for prediction.

Results

The best-performing tabular models (XGBoost, Random Forest, Gradient Boosting) achieved 93.3% accuracy on the original dataset; after hyperparameter tuning on the augmented dataset (20,000+ samples), XGBoost reached 91.2% accuracy distinguishing HαT, MCAS, and normal patients.

GAN-based augmentation failed statistical validation (KS testing showed distributions did not match the real data), while traditional noise-based augmentation preserved the original dataset's patterns — an important negative result showing not all synthetic data techniques generalize well to small medical datasets.

The microbiome pipeline identified specific bacteria (including Klebsiella pneumoniae strains and Ruminococcus obeum) with significantly different abundances between synthetic i-MCAS and normal samples, with two independent clustering methods (K-Means and GMM) converging on the same groupings — suggesting the approach is picking up genuine structure, not clustering artifacts.

The project's discussion emphasizes its proof-of-concept nature: the goal was to build and validate a reusable ML framework for MCD research, not to produce a clinically deployable diagnostic tool.

Skills

ML & Statistical Models
  • Logistic Regression (One-vs-Rest and Multinomial)
  • Decision Tree
  • Random Forest
  • Gradient Boosting
  • XGBoost
  • Support Vector Machines (SVM)
  • Multilayer Perceptron (MLP) / Neural Networks
  • K-Means Clustering
  • Gaussian Mixture Models (GMM)
  • Principal Component Analysis (PCA)
  • Generative Adversarial Networks (GANs)
  • Bootstrapping (resampling for synthetic data)
Model Development & Evaluation
  • Multiclass classification
  • Supervised & unsupervised learning
  • Hyperparameter tuning / GridSearchCV
  • Cross-validation
  • Stratified shuffle split
  • Class imbalance handling (SMOTE, SMOTEENN)
  • Confusion matrix analysis
  • Evaluation metrics (accuracy, precision, recall, F1)
  • Feature importance / interpretability
  • Explainable AI (xAI) principles
Data Processing & Feature Engineering
  • Exploratory Data Analysis (EDA)
  • Feature engineering
  • Domain-informed missing data imputation
  • One-hot encoding, label encoding
  • StandardScaler / MinMaxScaler normalization
  • Dimensionality reduction
  • Correlation analysis
  • Outlier detection
Synthetic Data & Statistical Validation
  • Data augmentation (noise-based + GAN-based)
  • Synthetic data generation and validation
  • Hypothesis testing (t-tests, Kolmogorov-Smirnov)
  • Distribution comparison/validation methods
Tools & Frameworks
  • Python, Jupyter Notebook
  • scikit-learn
  • TensorFlow / Keras
  • PyTorch
  • pandas, NumPy
  • Git/GitHub
  • joblib (model serialization)
  • GPU computing
Domain-Specific / Bioinformatics
  • Microbiome data analysis
  • 16S rRNA sequencing pipeline knowledge
  • Clinical/pathological tabular data analysis
  • Biomarker discovery
  • Reproducible research pipeline design

My Recent Projects

lines

Full Stack Web Development

Shortest Flight Path Algorithm

California Schools Data Visualization

Online Board Game Back-End

Java GUI

Trivia Quiz Website

Trivia Quiz Website

ML Kaggle Cirrhosis Data

ML Kaggle Cirrhosis Data

Mosaic Game Logical Agents

Mosaic Game Logical Agents

Data Mining ML

Data Mining ML

Relevant Courses

lines

Web Development: Full Stack

Spring 2024

Reinforced foundational programming skills while encouraging creativity, building upon the initial programming coursework.

HTML/CSSJavaScriptNode.jsExpressMongoDBVue.jsDatabase ManagementFull Stack Dev

Artificial Intelligence Practice

Spring 2024

Focused on the practical design and implementation of AI, with techniques in reasoning, planning, and learning. Students implemented AI concepts in software and learned how to evaluate their effectiveness.

Machine LearningPythonSearch AlgorithmsLogical AgentsAI

Information Visualization

Spring 2024

Introduced the design and development of visual representations for data exploration and analysis. Focused on principles of visual design, interaction, and evaluation, with practical assignments to reinforce skills.

D3.jsData VisualizationUX DesignDashboardsJavaScriptReact

Knowledge Discovery, Data Mining, and Machine Learning

Spring 2024

Covered automated data collection and analysis techniques for large-scale databases. Topics included model selection, tree methods, neural networks, and classification, with an emphasis on practical programming applications.

Data MiningModel TrainingMachine Learning AlgorithmsBig DataPythonJupyter Notebooks

Artificial Intelligence Principles

Fall 2023

Covered the foundational concepts of AI, including logical reasoning, uncertainty, and machine learning. Explored AI philosophies, problem-solving with search, and addressed philosophical issues within AI.

Machine LearningNeural NetworksAlgorithms

Software Engineering Principles

Fall 2023

Introduced key software engineering concepts such as development methodologies, high-level specifications, project management, and quality assurance, with an emphasis on ethical and sustainable practices. No programming was required.

AgileScrumProduct LifecycleUMLRequirements Gathering

Critical Systems Engineering

Fall 2023

Explored techniques for developing dependable socio-technical systems. Emphasized understanding system dependability and applying specialized software engineering techniques for reliable operation.

System DesignReliability EngineeringSafety Analysis

Java: Object-Oriented Programming, Modelling, and Design

Fall 2023

Strengthened skills in object-oriented design and implementation, essential for advanced programming assignments. Assumed prior programming experience equivalent to a Computer Science degree.

JavaOOPDesign PatternsBack-End

Computing Foundations: Data

Spring 2023

Explored fundamental computing principles with a focus on data structures, algorithms, and database management.

Data StructuresData ScienceSQLCS Fundamentals

Contact

lines

Lea Brody-Heine
lea_brody-heine@alumni.brown.edu

Education:
  • University of St Andrews
    MSc Computer Science
  • Brown University
    BA International and Public Affairs