Four examples of how AI, machine learning and data innovation are reshaping evaluation and research
By
Advancements in artificial intelligence, machine learning, and data collection technologies are rapidly transforming the landscape of research and evaluation, expanding evaluators’ methodological toolbox, and allowing them to ask and answer questions that would not have been feasible to answer before.
During the Global Impact Evaluation Forum 2025 in Rome, I had the pleasure of organizing a panel titled How will impact evaluation use AI, machine learning and innovations in data? The panel brought together researchers and evaluators who are actively experimenting with these technologies and applying them in their research and work.
Here are four examples of how AI, machine learning, and data innovation are reshaping research and evaluation:
1. Exploring natural processing language for climate adaptation in Bangladesh
Prabhmeet Kaur Matta, PhD student currently at the School of Computing, Information and Data Sciences at the University of California, San Diego, collected qualitative interviews as part of a quantitative survey in Bangladesh and analysed the interviews using Natural Language Processing (NLP).
Together with Rocco Zizzamia (University of Oxford) and Anindita Bhattacharjee (BRACT Institute of Governance and Development), Prabhmeet conducted a study to understand how households adapt to repeated climate shocks and explore the role of belief and expectations in Bangladesh.
As part of the study, they conducted survey interviews with 860 households which included open-ended questions where respondents freely shared their experiences on climate-shock responses.
Resulting in 645 hours of recorded audio, the team used AI to automatically transcribe and translate the open-ended responses and then analysed them with NLP tools such as topic modelling and sentiment analysis. The data revealed patterns in qualitative data that quantitative data may not have been able to capture. The presentation did not shy away from some of the challenges and adaptations the research team made during the process.
Prabhmeet showed how AI has the potential to enable data sources such as open-ended questions in structured survey interviews which can then be analysed using NLP tools. Compared with previous attempts made in this area, AI now shows significant innovation for languages, making these processes far more feasible and scalable than in the past. Yet Prabhmeet reminded us that human reviews remain key.
2. Creating a conversational platform for rapid experiments in Uganda
Stephan Dietrich, Assistant Professor at UNU-MERIT Maastricht University presented his work on developing a conversational platform for knowledge exchange at scale applied within the Kampala Capital City Authority.
Stephan’s team has developed a Large Language Model enabled platform in partnership with the Kampala Capital City Authority to communicate with constituencies in near realtime. The platform transcribes, translates and classifies incoming messages to generate immediate responses with the LLMs which runs locally to protect sensitive data.
Stephan highlighted the potential these platforms hold for A/B testing experiments and nudges in the interaction between public administration and the public. He also indicated these platforms’ future role in complementing traditional surveys to gather follow-up data.
3. Using local LLMs for transcript retrieval in a corporate level evaluation
Hannah Den Boer, Associate Evaluation Officer at IFAD’s Independent Office of Evaluation used local LLMs as a tool for enquiring transcript interview data as part of a corporate-level evaluation at IFAD.
Building on a guidance note for the thoughtful integration of AI for evaluation, Hannah shared how evaluators used a chatbot to query over 90 onehour interview transcripts of a corporatelevel evaluation conducted by the Office of Evaluation at IFAD.
While this AI solution had the advantage of significantly speeding up the analysis phase, evaluators had to factor in important considerations while engaging with the tool, including:
- Being aware of the risk of confirmation bias coming from LLMs, which the evaluators tried mitigating by constantly engaging with interactive prompts to explicitly request counterevidence to the claims made.
- Reducing the risk of hallucinations and incorrect responses, by prompting the system to return verbatim quotes, time staps and links to the original transcripts so they could manually verify the entries
- Noting the risks of overreliance on extracts and the implicit uniform weighting given to all the sources which the evaluators had to factor in while interpreting the findings
Hannah’s presentation showed how AI and LLMs have the potential to enhance efficiency in the evaluation process. However, a thoughtful approach to integrating such tools is required to maintain rigor.
4. Deploying computer vision for dietary assessment and nudging in Ghana and Vietnam
Aulo Gelli, Senior Research Fellow at the International Food Policy Research Institute (IFPRI) shared his experience in developing, piloting and implementing PlantVillage FRANI, an AI-based mobile phone-based application for dietary assessments, a collaboration with Penn State University’s PlantVillage, the University of Ghana and the National Institute of Nutrition in Vietnam.
FRANI uses smartphone photos with computer vision technology to identify foods and portion sizes from plate images. It then, produces nutrient estimates comparable to weighed records, and on par with 24hour recalls by trained dietitians, at a fraction of the cost.
Aulo presented studies in Ghana and Vietnam where data collection was combined with behavioural nudges to shift choices toward healthier dietary options.
Aulo’s presentation showed that while traditional dietary surveys are expensive, PlantVillage FRANI enables lower-cost high frequency monitoring data which were not feasible before.
Taking a system-wide governance view of AI
Nasim Motalebi, WFP Artificial Intelligence Lead, concluded the discussion reflecting on the governance systems and structure required to deploy AI solutions. She emphasized considerations regarding data privacy, procurement and legal frameworks, reminding the audience how adequate resourcing secure the “boring but vital” to ensure data ownership, integration, auditability and AI literacy across legal, ICT and evaluation functions. Finally, she focused on the business models and financing structures that reward longrun capability building over quick wins.
The panel presented four concrete examples of how researchers and evaluators are using AI, ranging from new data sources, such as images, unstructured text and audio, to novel methods for analysing them. Together, these examples demonstrated how AI can make processes that were previously possible but impractical, due to time or budget constraints, feasible and scalable.
I particularly appreciated the speakers’ openness in acknowledging that getting these tools “right” is far from straightforward, and that unforeseen challenges along the way demand ongoing flexibility and adaptation. The panellists also emphasized that alongside opportunities come complex significant challenges, including ethical considerations, data security risks, and issues of reproducibility.
As the field continues to navigate a steep learning curve, integrating these innovations into established research and evaluation practices requires vision, adaptability and a mindful critical approach.
