Francesco Massafra

Research Engineer

I build and evaluate AI systems — reinforcement learning, agentic planning, latent world models, and LLM evaluation — and I'm increasingly focused on efficient inference: making models faster and cheaper to serve. Reproducible experiments, clean research infrastructure.

Research

M.Sc. Artificial Intelligence student at the University of Amsterdam working as a research engineer: I design experiment pipelines, probe and evaluate models, and diagnose failures in reinforcement-learning and LLM systems. Two-time post-graduation research grant recipient; first author on arXiv:2603.05572.

Hi-LeWM: Hierarchical Planning in LeWorldModel

2026
Accepted @ WM@Booth — Workshop on World Models, Chicago Booth

Designed and evaluated a transformer-based hierarchical planner for goal-conditioned control, combining latent macro-actions, CEM/MPC planning, and receding-horizon execution over a frozen world model. Diagnosed off-manifold subgoal failures in long-horizon planning and improved robustness by constraining high-level search to data-supported macro-actions.

Code →

Open-Weight LLM Safety and Dataset Filtering

2026

Built reproducible LLM evaluation pipelines for prompt-based harmful-content classification across open-weight models, moderation datasets, and web-scale corpora. Implemented deterministic data mining, preprocessing, model-evaluation, and analysis workflows across Common Crawl, C4, FineWeb, and safety benchmarks.

Code →

Do Video Foundation Models Understand Intuitive Physics?

2026

Probed frozen V-JEPA, VideoMAE, and LTX-Video representations on IntPhys2 and MVP to evaluate whether pretrained video models encode intuitive-physics structure. Trained and compared linear, MLP, and temporal-attention readouts; analyzed layerwise emergence, temporal-control failures, and shortcut-driven benchmark behavior.

arXiv:2606.09646 → Code →

Machine Learning for analysis of Multiple Sclerosis cross-tissue bulk and single-cell transcriptomics data

2026

First author. End-to-end ML pipeline for MS transcriptomics across bulk microarray and single-cell RNA-seq (CD4+ and B-cells): XGBoost classification, SHAP feature attribution, and differential expression analysis. Models reached AUC 0.94 on CSF B-cells; SHAP gene selection proved complementary to classical DEA, surfacing novel mechanistic hypotheses and candidate biomarkers.

arXiv:2603.05572 →

Experience

Research Assistant — MEDICA Project, University of Pisa

2024 – 2025

Two-time post-graduation research grant recipient · Pisa, Italy

Built reproducible ML pipelines for transcriptomics classification, feature attribution, statistical analysis, and cross-dataset integration. Resulting work published as arXiv:2603.05572.

Full-Stack Developer — Wylit

2021 – 2025

Part-time · Remote / Hybrid

Built and owned full-stack production web platforms for 15+ client projects using React, Node.js, PostgreSQL, Docker, and Linux deployments, improving responsiveness and reducing load times by up to 40%.

Education

M.Sc. Artificial Intelligence — University of Amsterdam

2025 – present

GPA: 8.75/10

Relevant coursework: Reinforcement Learning, Deep Learning, NLP, Computer Vision, Information Retrieval.

B.Sc. Computer Science — University of Pisa

2021 – 2024

Final grade: 110/110

Thesis on ML for Multiple Sclerosis biomarker discovery.

Skills

Efficient Inference / Model Serving Reinforcement Learning Agentic Planning Latent World Models LLM Evaluation Model Probing Failure Analysis Reproducible Experiments Python C PyTorch pandas / NumPy / scikit-learn XGBoost / SHAP Docker Linux Git PostgreSQL

Contact

Interested in research, engineering, efficient inference, edge-AI deployment — or just a chat? Reach out.