AI evaluations & model behaviour
Design and lead evaluations of frontier AI models, translating ambiguous questions about model behaviour into measurable experiments. See our Science paper on AI evaluations for an example of this work.
AI Evaluations · Model Monitoring · Applied Machine Learning
At the European Commission AI Office, I work on emerging risks from general-purpose AI, including harmful manipulation, child safety and post market monitoring. I've been working on AI for +8 years, and my broader research and work experience spans the whole AI life cycle, from data quality to deployment and monitoring and policy. I also lead multidisciplinary technical teams and external research collaborations, and have published in Science, Nature Machine Intelligence, NeurIPS, AAAI, AIES and TMLR.
Before joining the European Commission, I was a Principal Investigator and Applied Skills Advisor at
a Marie Skłodowska-Curie Research Fellow
,
a statistician at the European Central Bank
,
and a consultant at Deloitte Robotics
.
I also held research positions and visits at
,
Barcelona Supercomputing Center
,
IIIA-CSIC
,
and
.
Before my research career, I competed as an elite swimmer
.
Design and lead evaluations of frontier AI models, translating ambiguous questions about model behaviour into measurable experiments. See our Science paper on AI evaluations for an example of this work.
Research emerging behaviours and risks from frontier AI, including post-market monitoring, manipulation, persuasion, over-reliance, child safety and mental-health interactions.
Lead multidisciplinary technical teams and external research collaborations on post-deployment evaluation and monitoring, building on my PhD research on detecting changes in model behaviour after deployment.
I've worked across research, technical projects and public institutions, combining applied research with product.
Design and lead technical evaluations of general-purpose AI models, including experimental work on manipulation, persuasion, child safety and post-deployment monitoring. Develop evaluation methodologies, coordinate external technical research teams, and monitoring emerging behaviours after deployment.
Technical and research advisory work for public, nonprofit and industry partners such as: British Antarctic Survey, WWF, Amnesty International, Roche, Astra Zeneca, Defence and Security Gov, Environmental Investigation Agency, John Muir Trust, CEFAS, Telenor, Turkcell, ...
Research on model monitoring, explainability and fairness within the EU-funded NoBias Innovative Training Network, including research visits at SCHUFA, the University of Pisa and BBC DataLab.
Statistical quality monitoring across EU Member States and applied machine-learning work for the European System of Central Banks.
Applied OCR and robotic process automation to audit and banking workflows, and trained Deloitte teams on UiPath.
My research spans the full AI lifecycle: data collection, data quality, preprocessing, modeling, and monitoring. I have led publications to NeurIPS, AAAI, AIES, TMLR and Science.