AI evaluations & model behaviour
Design and lead evaluations of frontier AI models, translating ambiguous questions about model behaviour into measurable experiments. See our Science paper on AI evaluations for an example of this work.
AI Evaluations · Model Behaviour · Machine Learning
At the European Commission AI Office, I lead technical work on emerging risks from general-purpose AI, including harmful manipulation, child safety and post market monitoring. My broader research and work experience spans the whole AI life cycle, from data quality to deployment and monitoring.
Before joining the European Commission, I was a Principal Investigator and Applied Skills Advisor at
,
a Marie Skłodowska-Curie Research Fellow within the NoBias network
,
a statistician at the European Central Bank
,
and a consultant at Deloitte Robotics
.
I also held research positions and visits at
,
Barcelona Supercomputing Center
,
IIIA-CSIC
,
and SCHUFA
.
Before my research career, I competed as an elite swimmer
.
Design and lead evaluations of frontier AI models, translating ambiguous questions about model behaviour into measurable experiments. See our Science paper on AI evaluations for an example of this work.
Research emerging behaviours and risks from frontier AI, including post-market monitoring, manipulation, persuasion, over-reliance, child safety and mental-health interactions.
Lead multidisciplinary technical teams and external research collaborations on post-deployment evaluation and monitoring, building on my PhD research on detecting changes in model behaviour after deployment.
I've worked across research, technical projects and public institutions, combining applied research with product.
Technical implementation and enforcement preparation of the AI Act for general-purpose AI. Lead work on post market monitoring, harmful manipulation and child safety, develop evaluations, coordinate external research collaborations and contribute to enforcement coordination with the Digital Services Act.
Technical and research advisory work for public, nonprofit and industry partners such as: British Antarctic Survey, WWF, Amnesty International, Roche, Astra Zeneca, Defence and Security Gov, Environmental Investigation Agency, John Muir Trust, CEFAS, Telenor, Turkcell, ...
Research on model monitoring, explainability and fairness within the EU-funded NoBias Innovative Training Network, including research visits at SCHUFA, the University of Pisa and BBC DataLab.
Statistical quality monitoring across EU Member States and applied machine-learning work for the European System of Central Banks.
Applied OCR and robotic process automation to audit and banking workflows, and trained Deloitte teams on UiPath.
My research spans the full AI lifecycle: data collection, data quality, preprocessing, modeling, and monitoring. I have led publications to NeurIPS, AAAI, AIES, TMLR and Science.