Yohan Zytoon

AI systems · ML engineering · Data science

Montréal, Québec, Canada

B.Sc. Computer Science · UdeM Graduating December 2026 Python · PyTorch · MLflow

I build AI systems that can be measured, debugged, and trusted—from agent workflows and evaluation tooling to the data pipelines underneath them.

My work sits between modeling and infrastructure: designing experiments, tracing failures across a stack, and turning useful prototypes into reproducible systems.

Data Scientist @ Intact Building VeroLoop

Measured impact

0.78 macro-F1

Root-cause analysis across prompts, retrieval, and workflows helped move a POC from 0.32 to 0.78.

How I work

Treat the model, data, evaluation, and deployment path as one system.

LLM agents Evaluation ML pipelines Reliability
Portrait of Yohan Zytoon

Current focus

Provider-neutral AI evaluation Agent reliability Budgeted search
Introduction

About

I care about the engineering around a model just as much as the model itself.

I’m a Computer Science undergraduate at the Université de Montréal working across applied AI, machine learning, and the infrastructure that makes both useful in production.

I’m most at home on problems that are slightly messy: an agent that is inconsistent, a pipeline that is too slow, or an experiment that cannot be reproduced. I like finding the failure mode, building the right instrumentation, and leaving behind a system other people can actually operate.

That has meant building stateful LLM workflows over enterprise data, evaluation harnesses that track every prompt and configuration, and ML pipelines that turn experimentation into repeatable delivery.

What I bring

A practical, systems-minded ML profile

  • Production-minded AI prototyping
  • Evaluation and failure analysis
  • Reproducible ML workflows
  • Clear communication across technical and business teams

Problems I’m drawn to

AI evaluation Agent systems ML infrastructure Applied modeling
Current Focus

Now

The systems and research questions I’m actively working on.

  • Building VeroLoop, an open-source evaluation layer for reproducible cross-provider AI benchmarks.
  • Exploring how to measure agent correctness, stability, latency, and cost without tying evaluation to one provider.
  • Studying budgeted combinatorial search with GFlowNets, Gaussian Processes, and active learning.
  • Developing stronger patterns for observable, versioned, and maintainable AI systems.
Academic Background

Education

Formal training that underpins my machine learning and computational work.

Université de Montréal

B.Sc. in Computer Science

Montréal, Québec · Expected Dec 2026

  • Machine Learning, Stochastic Processes, Probability & Statistical Inference, Optimization, Algorithm Design, and Distributed Systems.

Technical foundation

Algorithms, systems, statistics, and optimization

Built for work across the ML lifecycle

  • A foundation that connects model behavior to software constraints, data quality, and experimental design.
Professional Journey

Experience

Building reliable AI and data systems inside real operating environments.

Data Scientist I — DataLab, Adjuster Assist
Intact Financial Corporation
Montréal, Québec · May 2026 – Aug 2026
  • Built production LLM agents with LangChain and LangGraph, combining RAG, tool-calling, and stateful workflows over enterprise claims data.
  • Created an MLflow evaluation framework that logs prompts, workflows, configurations, metrics, and baseline deltas.
  • Improved fully stable predictions from 71% to 82% and overall stability by 5% through repeated-run evaluation and AI-generated diagnostics.
  • Traced failures across prompts, retrieval, and workflows, helping raise macro-F1 from 0.32 to 0.78.
  • Built a versioned, agent-accessible knowledge system with automated documentation and changelog updates.
Data Scientist Intern — Advanced Analytics
Desjardins Assurances générales (DAG)
Montréal, Québec · Sep 2025 – Dec 2025
  • Cut batch processing time by 80% through parallelized ingestion, better intermediate structures, and reusable transformations.
  • Engineered YAML-configured ML pipelines and integrated feature-store outputs into Azure ML scoring artifacts.
  • Built segmented correlation analysis to guide feature selection across subpopulations.
Business Intelligence Intern
Bombardier Recreational Products (BRP)
Montréal, Québec · May 2025 – Aug 2025
  • Designed Snowflake data models with automated reconciliation checks.
  • Built Power BI dashboards adopted across multiple business units.
Independent Work

Projects

Two projects that reflect how I think: build the infrastructure, define the measurement, then run the experiment.

VeroLoop

Python · Pydantic · MLflow · AsyncIO

A provider-neutral platform for evaluating AI systems.

  • Define a structured task once and benchmark it across OpenAI, Anthropic, and Gemini through one execution layer.
  • Typed contracts, provider adapters, schema compatibility, async concurrency, retries, and normalized errors.
  • Evaluates correctness, schema validity, stability, latency, and cost with versioned console, JSON, HTML, and MLflow reports.

Budgeted Combinatorial Search

PyTorch · GFlowNets · BoTorch · GPyTorch

Exploring an approximately 8-billion-state space under a 1,000-query oracle budget.

  • Compared Gaussian Process active learning, direct GFlowNet training, and a hybrid GFlowNet–AL pipeline.
  • Combined GP + UCB acquisition, diversity-aware selection, warm starts, delayed activation, and plausibility rewards.
  • Across five seeds, classical active learning reached a 3.7× higher valid-generation rate than direct GFlowNet training.
Toolkit

Skills

What I’m comfortable using day-to-day, grouped by how I actually think about my toolbox.

Programming & ML

  • Python, C/C++, SQL, Bash/Shell, R
  • PyTorch, scikit-learn, JAX, XGBoost
  • BoTorch, GPyTorch, Hugging Face
  • GFlowNets and statistical modeling

LLMs & Agents

  • LangChain and LangGraph
  • Azure OpenAI and RAG
  • Tool-calling and stateful workflows
  • Prompt engineering and evaluation

MLOps & Engineering

  • MLflow, Azure ML, WandB
  • Docker, Git/GitHub, Linux
  • YAML/Hydra and REST APIs
  • Evaluation, observability, reproducibility

Data & Infrastructure

  • Databricks and Snowflake
  • AWS: S3, SageMaker, Bedrock
  • Azure cloud services
  • PostgreSQL and MySQL
Foundation

Selected Coursework

The academic foundation behind my ML and systems work.

Machine Learning Probability & Statistical Inference Stochastic Processes Algorithm Design Optimization Distributed Systems
For Hiring Teams

For Teams & Recruiters

Where I can contribute on an AI, ML platform, or data science team.

  • AI systems: I can turn an LLM prototype into a typed, observable workflow with meaningful evaluation.
  • ML infrastructure: I build reproducible pipelines, evaluation harnesses, and internal tooling that reduce friction for a team.
  • Applied modeling: I’m comfortable moving between statistical analysis, model development, and controlled experiments.
  • Failure analysis: I trace errors through data, retrieval, prompts, and orchestration instead of treating a score as a black box.

If your team is working on reliable AI products, model evaluation, or the infrastructure around applied ML, I’d be glad to compare notes.

Connect

Contact

Let’s talk

I’m based in Montréal and open to roles and collaborations in AI engineering, ML infrastructure, machine learning, and data science.

If you think my profile fits your team, I’d be happy to connect and explore ideas.