Evidence before opinion
Every conclusion should be supported by observable evidence rather than assumptions or personal preference.
Hammed Babatunde Buari
Helping organisations improve AI through structured human judgement, evidence-based evaluation, and clear written feedback.
I evaluate AI-generated responses, agents, datasets, and digital workflows using structured quality frameworks—improving factual reliability, reasoning quality, instruction adherence, usability, and trust in AI systems.
Evaluation Purpose
Every day, people rely on AI to answer questions, interpret information, generate ideas, review documents, and support decisions. As these systems become more capable, the quality of their outputs becomes increasingly important.
Reliable AI is not achieved through model capability alone. It is strengthened through thoughtful human evaluation. By reviewing responses for factual accuracy, reasoning quality, instruction adherence, safety, clarity, and usability, evaluators help transform capable models into trustworthy systems.
My work focuses on bringing structured human judgement into that process through evidence-based review, consistent evaluation frameworks, and actionable written feedback.
Trustworthy AI begins with thoughtful evaluation.
Evaluation Philosophy
Every evaluation is an opportunity to improve trust in AI systems. My approach is guided by consistent principles that prioritise evidence, fairness, and meaningful human judgement.
Every conclusion should be supported by observable evidence rather than assumptions or personal preference.
Confident answers only create value when they are factually reliable and appropriately supported.
Applying structured evaluation standards ensures fairness, repeatability, and dependable quality decisions.
AI generates outputs. Human reviewers determine whether those outputs deserve trust.
Constructive evaluation transforms individual responses into stronger future AI performance.
Evidence of Evaluation
A growing collection of structured AI evaluations demonstrating evidence-based reasoning, rubric-led assessment, and actionable written feedback.
Response Review
Reviewing factual accuracy, reasoning quality, instruction adherence, completeness, safety, clarity, and usefulness.
Prompt Evaluation
Examining task clarity, missing constraints, expected outputs, evaluation risks, and opportunities to improve prompt quality.
Quality Rubrics
Applying structured criteria, scoring dimensions, written evidence, and clear decision rules across evaluation tasks.
Evaluation Walkthrough
A structured evaluation begins by understanding the task, examining the response against clear criteria, identifying evidence, and translating that analysis into an actionable decision.
Understand the task
Clarify what the response was expected to achieve, including required format, scope, audience, and completion criteria.
Inspect the response
Review the answer carefully to identify factual claims, assumptions, logical steps, omissions, and user-facing issues.
Apply the rubric
Evaluate accuracy, reasoning, instruction fit, safety, completeness, usability, and presentation against defined standards.
Record the evidence
Document the specific wording, omission, contradiction, or unsupported claim that justifies each evaluation decision.
Deliver the verdict
Summarise the result, explain its impact on the user, and recommend precise improvements for stronger future outputs.
Evaluation outcome
The final decision should be understandable to another reviewer, repeatable under the same rubric, and useful for improving the system.
Evaluation Case Study
This demonstration shows how I assess an AI-generated response using a structured rubric, traceable evidence, calibrated scoring, and actionable written feedback.
01 · Evaluation Task
Explain why regular exercise is beneficial for adults. Provide three clear benefits in fewer than 120 words and avoid making unsupported medical claims.
02 · AI Response
Regular exercise is essential because it prevents nearly every major disease and guarantees a longer life. First, it completely eliminates the risk of heart disease by strengthening the heart. Second, exercise permanently cures anxiety and depression because physical activity releases positive chemicals in the brain. Third, exercising every day rapidly increases energy, improves sleep, strengthens immunity, prevents ageing, and ensures that adults remain healthy throughout life.
03 · Evidence Review
“prevents nearly every major disease and guarantees a longer life”
The wording overstates the evidence and presents health outcomes as guaranteed rather than describing exercise as reducing risk.
“completely eliminates the risk of heart disease”
Exercise can support cardiovascular health, but it does not eliminate all risk of heart disease.
“permanently cures anxiety and depression”
Physical activity may support mental wellbeing, but it should not be presented as a permanent cure for clinical conditions.
“rapidly increases energy, improves sleep, strengthens immunity...”
The response lists several additional benefits without clearly organising the answer into the requested three distinct points.
04 · Rubric Assessment
Scale: 1 = poor · 5 = excellent
05 · Final Verdict
The answer is understandable but fails the factual reliability and safety requirements because it presents health outcomes as guaranteed facts.
Replace absolute language such as “guarantees,” “completely eliminates,” and “permanently cures” with appropriately qualified statements. Present exactly three distinct benefits, explain each clearly, and describe exercise as supporting health or reducing risk rather than guaranteeing medical outcomes.
06 · Improved Response
Regular exercise offers several important benefits for adults. First, it supports cardiovascular health and can help reduce the risk of conditions such as heart disease. Second, physical activity can improve mood and reduce stress by supporting mental wellbeing. Third, regular movement helps maintain muscle strength, mobility, and healthy energy levels, which can make everyday activities easier. The most suitable routine depends on a person’s health, ability, and circumstances, so gradual and consistent activity is generally more helpful than an extreme approach.
Evaluation Practice
Evaluation is not a final inspection added after the work is complete. It is a structured practice for understanding intent, examining evidence, identifying meaningful weaknesses, and explaining how an output can become more accurate, useful, and trustworthy.
Review Method
Define the task, its purpose, intended audience, constraints, and required standard before making any judgement.
Translate the objective into clear criteria covering correctness, relevance, completeness, clarity, usability, and presentation.
Examine the output closely, separate observable evidence from assumption, and test whether each conclusion is supported.
Distinguish critical failures from minor imperfections and focus attention on issues that materially affect quality or trust.
Provide precise, actionable feedback that identifies the issue, explains its impact, and shows what a stronger result requires.
Practice Map
Verify factual claims, calculations, internal consistency, assumptions, and alignment with the available evidence.
Determine whether the response addresses the actual task directly, respects its constraints, and avoids unnecessary material.
Assess whether conclusions follow logically from the evidence and whether important assumptions, trade-offs, or uncertainties are clear.
Examine structure, language, hierarchy, and explanation to determine whether the output can be understood without avoidable effort.
Consider whether the result helps its intended user make a decision, complete a task, or move confidently to the next step.
Review visual hierarchy, consistency, accessibility, and polish where presentation quality forms part of the work product.
My work combines AI evaluation, data validation, UX review, annotation, ranking, healthcare operations, and structured written feedback.
Assessing factual accuracy, reasoning quality, completeness, clarity, consistency, and instruction adherence.
Reviewing task completion, workflow quality, ambiguity, output reliability, and user-facing usefulness.
Validating datasets for completeness, duplication, anomalies, consistency, reporting quality, and operational readiness.
I evaluate outputs through practical dimensions that matter in real user-facing AI systems.
Is the response factually reliable, complete, and free from unsupported claims?
Does the answer follow a logical structure and avoid hidden assumptions or contradictions?
Does the output follow the user’s request, constraints, format, and intended task?
Does the output avoid harmful, biased, misleading, or overconfident content?
Is the response clear, concise, well-structured, and useful for the end user?
Does the output look polished, readable, correctly formatted, and professionally delivered?
Evaluation Archive
This archive publishes completed evaluation records with their evidence, comparative findings, final judgements, and documented limitations preserved as part of one accountable body of work.
Canonical Record Library
Each archive entry records the evaluation brief, assessed responses, comparative reasoning, final ordering, synthesis, and canonical record statement.
Every published evaluation receives a permanent AI-R identifier and remains preserved in its original editorial state. New evaluations are appended to this archive without altering previously published records.
Independent comparative evaluation examining reasoning quality, trade-off analysis, recommendation strength, and editorial judgement across three frontier AI systems.
Evidence-led heuristic and direct-interaction evaluation examining task clarity, transparency, data trust, validation behaviour, interpretation and user control.
Published evaluation records are preserved in their original editorial form. Where revisions, corrections, or subsequent analyses are required, they are issued as new canonical records rather than replacing previously published work.
Future evaluations will be released as AI-R-002, AI-R-003, and subsequent canonical records as new evidence-based assessments are completed. Only fully reviewed evaluations are admitted into the permanent archive.
Independent comparative evaluation demonstrating structured reasoning, evidence synthesis, comparative judgement and canonical record construction across multiple AI systems.
The following publication preserves the complete evaluation exactly as conducted. The original evaluation brief, assessed responses, independent analyses, comparative reasoning, editorial synthesis and final judgement are presented sequentially to maintain transparency, reproducibility and permanent archival integrity.
Evaluation Prompt
A local government has a budget to implement only one of the following initiatives this year: expanding mental health services, improving public transport, or upgrading flood defences. Explain the strongest arguments for each option, discuss the trade-offs, and recommend one option while acknowledging uncertainty and the circumstances under which a different decision would be justified.
This prompt was selected to evaluate each system's ability to balance competing public priorities, reason under uncertainty, acknowledge trade-offs, avoid false certainty, and produce a justified recommendation supported by transparent analysis.
AI System: ChatGPT
A local government choosing between these three initiatives faces a classic resource-allocation problem: each addresses an important public need, but they produce different types of benefits over different time horizons.
The strongest argument for expanding mental health services is that improved mental health can have wide-ranging effects across the community. Better access to counseling, crisis intervention, and early treatment may improve quality of life, reduce pressure on emergency services, support employment and education, and help prevent more serious health problems. If current services are difficult to access or have long waiting lists, additional investment could address an immediate unmet need.
The main trade-off is that many benefits are difficult to measure directly and may take time to become visible. Funding mental health services also does not directly reduce risks from physical infrastructure failures or improve transportation networks.
The strongest argument for improving public transport is that it benefits a broad cross-section of residents. Reliable, affordable transport can improve access to jobs, healthcare, schools, and local businesses while reducing congestion and emissions. It may also promote economic development by making the area more accessible.
The trade-off is that major transport improvements can require substantial ongoing operating and maintenance costs. Benefits may also be unevenly distributed if improvements primarily serve certain neighborhoods or travel corridors.
The strongest argument for upgrading flood defences is risk reduction. If the community faces a meaningful flood hazard, investments can prevent loss of life, protect homes and businesses, reduce future recovery costs, and increase community resilience. Preventing a single major flood event can avoid damages that far exceed the initial investment.
The trade-off is that if flooding is relatively rare or the current level of protection is already adequate, the investment may appear to deliver little immediate visible benefit compared with services residents use every day. Funds dedicated to flood defences also cannot simultaneously address pressing health or transportation needs.
If I had to recommend one option without additional local information, I would recommend upgrading flood defences, provided the community faces a credible and significant flood risk.
This recommendation rests on the principle that governments should prioritize preventing catastrophic harm when the likelihood and consequences are substantial. Flood events can result in deaths, displacement, extensive property damage, environmental impacts, and long-term economic disruption. Because these losses are often irreversible or extremely costly to recover from, preventive infrastructure can provide significant value even if floods occur infrequently.
This recommendation depends heavily on local circumstances. A different decision would be justified if:
Ultimately, the best decision depends on evidence such as local flood risk assessments, health service demand, transport performance, population needs, and cost-effectiveness analyses. When these data differ, the appropriate priority may differ as well. A transparent decision-making process should weigh both the expected benefits and the uncertainty surrounding each option.
Response R-01 demonstrates structured analytical reasoning throughout the evaluation. Each policy option is examined independently before a final recommendation is presented, maintaining a logical progression from description to conclusion.
Response R-01 satisfies the evaluation brief through balanced reasoning, appropriate acknowledgement of uncertainty, and a justified recommendation. The response demonstrates good analytical quality, while leaving scope for stronger comparative synthesis across the competing policy options.
AI System: Claude
Since no context about the specific town is given, I'll frame this generically and flag what local specifics would change the answer.
Strongest case: high and rising demand post-pandemic, chronic underfunding relative to physical health, and strong evidence that early intervention reduces downstream costs (A&E visits, police call-outs, homelessness, lost productivity). Benefits are often invisible until a crisis is prevented, which makes it politically harder to champion but ethically compelling — it protects the most vulnerable residents.
Strongest case: touches nearly everyone daily, has clear economic multiplier effects (access to jobs, reduced congestion), and contributes to climate goals. Improvements are highly visible, which builds public trust in local government. Downside: benefits tend to accrue unevenly (whoever lives near the improved routes) and capital costs can spiral.
Strongest case: this is the closest to a genuine public-safety insurance policy — the cost of inaction can be catastrophic and irreversible (loss of life, property destruction, insurance market collapse in affected areas). Unlike the other two, delay doesn't just mean slower progress; it can mean a specific disaster happening on your watch. This is the only option with a plausible “hard deadline” (before the next major storm season).
Absent other information, I'd lean toward flood defences — the asymmetry of outcomes (catastrophic, irreversible harm vs. everything else being serious but recoverable) usually justifies prioritizing risk reduction when the risk is credible and imminent.
But this recommendation flips under identifiable conditions:
The honest answer is that this decision hinges heavily on local risk data, current service conditions, and what other funding streams exist — a generic ranking is a reasonable starting prior, not a substitute for looking at your specific numbers.
Response R-02 provides a concise and analytically structured treatment of the three policy options. It distinguishes between ongoing quality-of-life benefits and protection against low-probability, high-impact harm before presenting a conditional recommendation.
Response R-02 demonstrates strong analytical structure, clear communication, and meaningful treatment of uncertainty. Its distinction between ordinary service benefits and catastrophic tail-risk protection provides a persuasive basis for the conditional recommendation. The response remains balanced by identifying clear conditions that would reverse its conclusion, although its comparative judgement would benefit from a more systematic decision framework and stronger qualification of unsupported general claims.
AI System: Gemini
Local governments face constant tension between immediate human suffering, structural economic efficiency, and existential community safety. Choosing between expanding mental health services, improving public transport, and upgrading flood defences requires weighing invisible, cumulative crises against acute physical risks.
Recommendation: Upgrading flood defences represents the most prudent baseline choice for a single-budget constraint, provided the local government faces any demonstrable hydrological risk.
The Justifying Logic: Infrastructure security is the ultimate prerequisite for municipal existence. A severe flood can instantly wipe out transport networks, overwhelm local health systems, and displace the very population mental health initiatives aim to serve. Physical safety and asset preservation must form the non-negotiable floor upon which all other social and economic progress is built.
When a Different Decision Would Be Justified:
Which of these three pressure points—social wellbeing, mobility, or physical safety—is currently presenting the most acute daily strain in your local context?
Response R-03 adopts a policy-oriented and strategically structured approach to the decision problem. It evaluates each initiative through human, economic and infrastructure perspectives before presenting a conditional recommendation grounded in municipal risk management.
Response R-03 is a mature, coherent and persuasive policy analysis. It balances competing public priorities effectively, communicates uncertainty appropriately and provides a defensible recommendation without overstating certainty. The principal opportunity for improvement is a more explicit evidence-based comparison framework and stronger linkage between conclusions and measurable decision criteria.
All three responses demonstrated coherent reasoning and addressed the central allocation problem directly. Response R-01 followed a balanced explanatory structure, Response R-02 introduced a stronger risk-based analytical framework through conditional reasoning, while Response R-03 adopted the broadest strategic perspective by linking human, economic and infrastructure considerations into a unified policy narrative.
Each response presented arguments supporting all three initiatives before reaching a recommendation. R-01 maintained consistent balance throughout, R-02 contrasted competing priorities through explicit policy trade-offs, and R-03 expanded the discussion by framing each initiative within wider municipal resilience, social wellbeing and economic development objectives.
Comparative differences became most apparent within trade-off reasoning. R-01 identified practical strengths and weaknesses for every initiative. R-02 distinguished between chronic service improvement and catastrophic tail-risk protection, providing a clear conceptual comparison. R-03 integrated operational, financial and strategic consequences into a broader municipal decision-making framework.
All three responses appropriately acknowledged uncertainty and avoided presenting their recommendations as universally applicable. Each identified conditions under which an alternative decision would be justified, demonstrating appropriate awareness of evidence-dependent public policy evaluation.
Each response ultimately recommended prioritising flood defences under credible flood-risk conditions. Differences arose primarily in the depth of supporting justification rather than the recommendation itself. R-03 provided the broadest policy rationale, R-02 emphasised asymmetry between catastrophic and recoverable outcomes, while R-01 focused on preventing irreversible harm through risk reduction.
All responses were logically organised and professionally presented. R-01 adopted a traditional explanatory format, R-02 emphasised concise analytical precision, and R-03 employed the most formal report-style structure through clearly separated arguments, trade-offs and recommendation sections.
Independent evaluation demonstrated that each system successfully satisfied the evaluation brief while exhibiting distinct analytical characteristics. The differences observed related primarily to depth of synthesis, treatment of competing priorities, and breadth of policy reasoning rather than factual correctness or instruction adherence. These comparative observations provide the evidential basis for the consolidated findings presented in the subsequent evaluation record.
The following matrix summarises the comparative findings presented in the preceding analysis. It serves as a structured reference and does not replace the detailed independent evaluations or comparative reasoning.
| Evaluation Dimension | Response R-01 | Response R-02 | Response R-03 |
|---|---|---|---|
| Instruction Adherence | Strong | Strong | Strong |
| Reasoning Quality | Strong | Very Strong | Excellent |
| Trade-off Analysis | Good | Very Strong | Excellent |
| Acknowledgement of Uncertainty | Strong | Excellent | Good |
| Recommendation Quality | Well Justified | Very Well Justified | Most Comprehensive |
| Communication Quality | Excellent | Excellent | Excellent |
| Overall Position | Third | Second | First |
The matrix provides a concise comparative reference only. Final editorial judgement is established through the Comparative Outcome and Canonical Record Statement that follow.
AI-R-001 evaluated three frontier AI systems against a single policy reasoning task requiring balanced analysis, acknowledgement of uncertainty, and a justified final recommendation. Each response was examined independently before comparative synthesis was performed to establish the canonical evaluation record.
The assessment considered reasoning quality, balance, structure, completeness, communication, and adherence to the evaluation brief. Comparative findings were then consolidated into one evidence-based record rather than isolated model reviews.
Following independent assessment and comparative synthesis, the evaluated responses were ranked according to their overall performance against the evaluation brief.
AI-R-001 establishes the first canonical evaluation record within this archive. Following independent assessment, comparative analysis, and record-level synthesis, the final comparative ordering for this evaluation is: Gemini (R-03), Claude (R-02), and ChatGPT (R-01). This ordering is specific to the evaluation brief and should not be interpreted as a universal ranking across unrelated tasks or domains.
Evaluate AI-generated responses, technical content, and enterprise deliverables against structured evaluation frameworks, assessing factual accuracy, reasoning quality, clarity, and instruction adherence while producing concise, evidence-based feedback.
Evaluation relevance: Rubric-led assessment · factual and reasoning review · instruction adherence · evidence-based feedback
Review datasets through annotation, classification, and quality-assurance workflows, identifying inconsistencies, bias, and data-quality issues while maintaining accuracy and consistency across evaluation tasks.
Evaluation relevance: Annotation quality · classification consistency · bias review · dataset quality assurance
Validate healthcare and public-health datasets for completeness, accuracy, consistency, and reporting readiness through systematic data-quality checks.
Evaluation relevance: Data completeness · accuracy validation · consistency checks · reporting readiness
Developed web applications and dashboards while strengthening testing, debugging, API integration, and documentation skills that support the evaluation of software and AI products.
Evaluation relevance: Software testing · debugging · API integration · technical documentation
Evaluated AI-generated responses against structured rubrics, identifying factual inaccuracies, reasoning gaps, and instruction-following issues.
Evidence demonstrated: Structured rubric application · factual accuracy review · reasoning assessment · instruction-following analysis
Conducted structured evaluations of digital products using heuristic analysis and evidence-based recommendations.
Evidence demonstrated: Heuristic evaluation · usability issue identification · structured product review · evidence-based recommendations
Designed practical evaluation frameworks and quality criteria for consistent, repeatable assessment.
Evidence demonstrated: Evaluation criteria design · repeatable assessment structure · quality standards · consistency-focused review logic
Built a documentation-first portfolio showcasing evaluation philosophy, methods, and case studies.
Evidence demonstrated: Documentation-first architecture · evaluation-method communication · case-study synthesis · structured portfolio presentation
Best aligned to AI Evaluator, AI Trainer, LLM Response Evaluator, Data Quality Analyst, Human Feedback Reviewer, Healthcare Operations Evaluator, UX Quality Reviewer, and Technical Assessment Contributor roles.