Are we headed for the Wild West of AI use in evaluation?
By Nea-Mari Heinonen
Ministry for Foreign Affairs of Finland

Artificial intelligence (AI) is increasingly presented as a solution to a wide range of organisational challenges in public administrations and multilateral organisations. Across sectors, there is growing pressure to adopt AI, often accompanied by assumptions that it will automatically improve efficiency. For the evaluation community, this raises new questions, such as: How do we prevent the emergence of a Wild West of AI use that can lead to pseudo-evaluation and pseudo-evidence? Figuratively, ‘the Wild West’ refers to “something that is chaotic and uncontrolled especially due to a lack of regulation or oversight”. Using this metaphor, this blog raises four key concerns. Coincidentally, around the time of publication, others have also found the metaphor apt for describing AI-related challenges.
There is no such thing as "The AI"
Firstly, ”AI” is often discussed as though it were a single technology or autonomous actor, despite encompassing a diverse range of human-designed tools, methods and techniques. In many discussions, AI has become shorthand for generative AI, and for specific chat-based tools such as Copilot or ChatGPT.
This lack of precision creates conceptual confusion. Different people understand AI differently, depending on the technologies available in their organisation and their own level of familiarity with them. When "AI" is used as a catch-all term, there is a risk of creating a Wild West of concepts in which people are discussing different things while believing they are talking about the same thing. This can obscure which tools are appropriate for specific tasks and data types, potentially leading to a Wild West of conceptual confusion. The evaluation community needs to become more specific about the technologies that are being discussed and the purposes for which they are used.
Who decides what appropriate use looks like?
The second concern relates to decision-making in organisations: How to avoid hasty decisions about AI use and manage unrealistic or misinformed expectations? Who decides what appropriate use is, and will the subject-matter experts be involved in such decisions? Many evaluation experts can envisage a situation where someone with limited evaluation knowledge proposes an apparently simple solution: "Why not ask a chatbot whether the project was successful?" While well-intentioned, such suggestions underestimate the nature of evaluation.
At present, many decisions about AI use are made at the level of individual users. People experiment with tools as part of their personal AI uptake, often without guidance, standards or oversight. At the same time, senior decision-makers in different organisations face pressure to increase efficiency. This can result in a fragmented environment in which AI application emerges through ad hoc decisions. Without appropriate governance, this creates a real risk of producing what might be called a Wild West of pseudo-evaluation: activities that resemble evaluation on the surface but lack the methodological, quality and other requirements to meet evaluation standards and principles. The cart is being put before the horse when AI tools are increasingly used to analyse information as method choices rather than purposeful complementary tools. As organisations move towards agentic AI systems, this challenge becomes even more complex. The evaluation community should not take the demand for rigorous evaluations for granted.
When accepted uses collide
Thirdly, even if evaluation professionals are successfully engaged to define appropriate uses of AI, another problem remains. Different organisations reach different conclusions about what is acceptable use. Internal regulations, policies and risk tolerances vary considerably. A practice considered appropriate by one organisation may be prohibited by another – for example, AI-assisted analysis of interview transcripts.
This creates practical challenges for evaluators, commissioners, and informants, potentially leading to a Wild West of accepted use. Informed consent will not solve this multi-layered problem alone, assuming all informants are able to make informed choices about such highly technical matters. Informed rejection, on the other hand, can lead to unequal representation in the data. Navigating the rules becomes increasingly difficult, and will require ongoing dialogue, transparency and collaboration across organisations. It will also require willingness to revert to traditional evaluation methods when necessary to ensure adequate representation.
Selection and differing uses generate pseudo-evidence
Perhaps the greatest risk concerns the long-term quality of evidence itself. Much of the public debate around AI has focused on hallucinations and bias. These are important concerns, but equally important is the issue of selection. When tools such as Copilot or Claude summarise information, they make choices about what to highlight, prioritise and omit. Many users assume these processes are objective and systematic. In reality, it is not possible to know exactly how information is selected, nor does the same task produce identical outputs over time.
Most users understand that AI-generated outputs should not be accepted uncritically but many overlook a subtler issue. Summaries can be factually correct while emphasising different aspects of the source material. A small experiment using three different large language models (LLMs) to summarise evaluation findings showed that all outputs were factually correct, yet they differed in formulation, emphasis, presentation and balance of content. One also generated recommendations, although recommendations were not part of the materials.
Summarising evaluation findings passages with different LLMs | |
Description: A set of 11 text passages on evaluation findings were extracted and tested using three different LLMs. Prompt: Summarise the content of the document in about 100 words. The results were cross-checked against the original passages. All were factually correct but differed in what became emphasized through selection and choice of words, framing, level of detail, and in balance of substantive elements, detectable through manual color-coding. One LLM also generated recommendations to maximise impact (in its opinion), although not requested nor was such content included in the original materials. Bold added by author to highlight some such features. | |
“[Country] is leveraging high‑level advocacy, strategic partnerships with tech giants and targeted financing to embed gender equality in global digital‑development policies and projects. While it has achieved notable policy wins (e.g., gender clause in the [name of compact]) and piloted innovative solutions that benefit women and vulnerable groups, inclusion outcomes remain uneven—especially for low‑income, women‑led micro‑enterprises. The main challenges are rigid financing tools, limited visibility of [country’s] solutions, and insufficient grassroots digital literacy. To maximise impact, [country] needs more flexible, harmonised funding mechanisms, stronger SME engagement, and intensified capacity‑building for the poorest and most marginalized users.” (LLM1) | |
“[Country] has strongly promoted inclusion, gender equality, and digital development through multilateral influence and bilateral projects. Its leadership in the [name of campaign] and co-leading of the [name of coalition] helped address the digital gender divide and online safety. [Country] contributed to shaping the [name of compact] and supported digital public goods, AI governance, and connectivity. Bilateral initiatives, such as digital land administration in [country] and innovation ecosystem support in Africa and [country], showed significant impact. However, challenges persist, including mixed inclusion outcomes, limited benefits for disadvantaged groups, difficult SME mobilisation, underfunded instruments, and reduced visibility in multilateral programmes.” (LLM2) | |
“[Country] has played a leading role in promoting inclusive digital development, particularly through multilateral engagement. It has influenced global processes such as the [name of compact] and co-led the [name of coalition], advancing gender equality, digital inclusion, and online safety. Bilateral initiatives, including digital land registries in [country], have supported women’s rights and access to finance. However, inclusion outcomes remain mixed, especially in entrepreneurship and job creation, with limited benefits for the most vulnerable groups. Challenges include underfinanced multilateral instruments, reduced visibility, limited SME engagement, and persistent global digital divides in connectivity, affordability, electricity access, and digital skills.” (LLM3) | |
Legend | [blue=description] [green=positive results] [yellow=mixed results] [red=challenges] [purple=LLM-generated recommendations] |
Analysis and summarisation are inherently processes of selection. Evaluators make such choices deliberately, guided by purpose, context and methodological judgement. AI systems also make selections, but often through processes that remain opaque to users. The effects are multiplied by the various types of accepted uses of AI.
The danger is not necessarily just incorrect information but the gradual erosion of nuance, context and complexity. As AI-generated summaries are used to create further summaries, over time, there is a risk creating a chain of "summaries of summaries" in which essential details are lost, causing the fourth concern, a Wild West of pseudo-evidence, through selection and the cumulative loss of context and meaning over time. This bears profound consequences, altering the very nature of evidence.
Towards thoughtful application
The use of AI technologies in evaluation raises important questions regarding their intended purposes, the boundaries of acceptable applications, and the implications of differing interpretations across organisations. While definitive answers may not yet exist, it is essential that the evaluation community begins to engage with these issues.
The evaluation community has much to gain from engaging with various applications of AI. As initiatives such as OpenEval.fi demonstrate, there are genuine opportunities to support evaluative work in new ways. However, the future will not be determined by the technology alone but also by the human ability to govern its use. The key question is whether evaluation professionals will have an active role in defining appropriate uses of AI within various organisations.
The evaluation community must also continue to raise awareness of "what evaluation is". AI applications may increasingly outperform humans at processing data. However, evaluation entails much more: interpreting significance, meaning and value; and answering evaluation questions that are complex, contested, value-laden, influenced by context, and open to multiple interpretations.
The ultimate challenge is not whether AI should be used in evaluation, but how it can be thoughtfully employed in ways that strengthen the quality, credibility and integrity of evidence on which decisions are based. Failure risks entering a Wild West of pseudo-evaluation and pseudo-evidence due to unthoughtful application of AI. Success can help establish AI as a valuable tool in service of rigorous evaluation, rather than a poor substitute for it.

Nea-Mari Heinonen (MA, MSSc) is the Lead Evaluation Specialist and Deputy Director of the Development Evaluation Unit of the Ministry for Foreign Affairs of Finland. With 20 years in international development, she has piloted AI and data science approaches and is currently pursuing a PhD in social data science at the University of Helsinki. Connect with Nea-Mari on LinkedIn.
The views expressed in this blog post are those of the author and do not necessarily reflect the views of her affiliated organisations.
Acknowledgement of the use of AI: Copilot was used to create the first draft of this blog post based on the author’s original panel speeches prepared for gLocal, June 2026. All content is original by the author.
Disclaimer: The content of the blog is the responsibility of the author(s) and does not necessarily reflect the views of Eval4Action co-leaders and partners.



Comments