Skip to content
Freelance Productivity
9 min readKemi Okoro

Multilingual AI research: check the baseline before claiming an update

In multilingual AI research, a sourced proposal need not prove an update. One Farsi workbook shows why the relevant portal baseline still matters.

Editorial panels show a proposed value beside a baseline marked not observed; their update comparison remains unresolved.

Fluent Farsi does not establish original-language evidence coverage. In a procurement experiment published on 3 October 2026, Roya Pakzad reports fluent Farsi answers alongside weaker Iran retrieval. Her US/English and Iran/Farsi tasks also varied jurisdiction and infrastructure, so this was not a language-only comparison. [1]

The sharper detail is in a linked workbook: cell D2 marks the current portal value not observed, while E2 offers Iran, Islamic Rep.. The row is included for completeness, not as a demonstrated correction. [1]

For multilingual AI research, I would keep a useful proposal at the level its evidence supports. Before calling it an update, establish what it is meant to replace. A defensible fact can still leave that comparison unanswered.

Editorial panels show a proposed value beside a baseline marked not observed; their update comparison remains unresolved.

Editorial analysis: a proposed value does not establish an update against an unobserved baseline.

TL;DR: A sourced proposal may be useful without proving an update. Verify the proposed fact if that is all the brief needs. A change claim also needs the relevant baseline. This workbook discloses its gap; Pakzad's experiment does not isolate a language effect.

The workbook leaves the current value unobserved

Pakzad asked agents to research country-profile updates in the World Bank's Global Public Procurement Database (GPPD). [1]

Open the published Iran_Claude.xlsx workbook. On its field sheet, D1 labels the current portal value and E1 the proposed value. In the country-name row, D2 begins with مشاهده‌نشده, paraphrased here as “not observed”. E2 contains Iran, Islamic Rep.. These are the contents of two different columns, not an old name and its confirmed replacement. [1]

Workbook reconstruction: D2 is not observed, E2 proposes Iran, Islamic Rep.; the country-name row is included for completeness.

_Reconstruction of Iran_Claude.xlsx: D2 begins with not observed, E2 proposes Iran, Islamic Rep., and I2 includes the row for completeness._ [1]

That Farsi label is an excerpt, not the entire cell. The rest of D2 says the portal page could not be loaded from the session. The English readings here are editorial paraphrases, not certified translations. I2 says the country-name row was included for completeness. It does not establish that this field was missing, outdated or incorrectly corrected. [1]

The methodology sheet makes the wider limit explicit. A2 says the current portal values were not observed and need direct page observation. [1] This is useful disclosure. The workbook has preserved the difference a reviewer needs, rather than silently presenting every proposal as a completed update.

Only one workbook was decoded for this article. Its country-name source and other proposed values were not independently verified. Reading the cells establishes what the artifact records; it does not authenticate every action behind those statements.

The distinction matters even if a reviewer later verifies the proposal. That would supply evidence for E2, while D2 would still lack an observed portal value. The portal could already contain the proposed name, contain another value, or leave the field empty. Those are conditional possibilities, not findings about the live page. A supported proposal alone cannot distinguish them.

A verified proposed value does not establish what changed when the baseline was not observed. That is the inference I draw from these columns. It identifies a missing comparison, without declaring the proposed fact false.

For a freelancer asked to find candidate information, the proposal and its source lead may be a useful deliverable. A commission to identify missing or stale portal entries needs evidence of those entries too. Do not turn “not observed” into “missing” or “fixed” during handoff. The next question is whether the substitute source supplies the comparison the portal could not.

An official fallback can change the comparison

The workbook's methodology cell A3 names the World Bank's GPPD data-collection report as the source of its field inventory. The linked report has 2020 page headers. [1] [2] It helps explain the structure of the spreadsheet. It does not supply an observation of the live portal's country-name entry.

Pakzad separately reports that Claude could not load the live country profiles for either Iran or the US and continued through the GPPD DataBank API. She describes that fallback as 2018 data, compared with a 2022 portal profile. [1] That chronology comes from her account; decoding this workbook does not independently reconstruct it. Neither year automatically means current in 2026.

Keep the dates attached to their objects. A collection report's date identifies a document. A dataset's reference year identifies the period of its data. The portal-profile year is another comparison point in Pakzad's account. An official publisher does not make those objects interchangeable.

The substitution can change what the task is able to answer. If the required portal cannot be reached, an alternate document may still yield useful information and let an agent return a spreadsheet with the requested columns. Yet the supported question can shift from which portal entries need correction to which candidate values can be assembled elsewhere. In this workbook, the unobserved-current-value column makes that shift visible. [1]

For a dated background summary, an older official source might be entirely appropriate. For an update, the reviewer needs to establish that the substitute covers the relevant field and period, and what relationship it has to the entry being revised. Otherwise, “official” describes the source's publisher while leaving the comparison unsettled.

That check can be performed on the documents now. Whether the agent originally reached a particular source is a separate question.

A missing citation does not prove non-access

There is positive evidence of an access problem in the published Claude Iran transcript. At about [00:50], it records connect_rejected for www.globalpublicprocurementdata.org:443, with this diagnostic: [1]

gateway answered 403 to CONNECT (policy denial or upstream failure)

The wording leaves the cause unresolved. It records a failed attempt in that published record, not every possible route or later attempt. It does not diagnose Iranian filtering.

The transcript is derived from screen recordings, not an authenticated execution log. Its preamble describes OCR extraction with visual checks and warns that brief changes may be missed and text may retain errors. This article did not independently inspect the original video. [1] The diagnostic is therefore evidence from a published visual transcript, not a replicated network measurement.

A citation has a different job: it points a reader towards a source. Opening that source and checking a supporting passage can establish claim support now, without proving how the agent originally obtained the answer. OpenAI's search guidance says responses may include citations and recommends checking whether a source supports the answer, as well as its date and authority. [3] An absent citation alone cannot establish non-access or identify training data as the answer's origin.

Pakzad also requested retrospective self-reports, which she says she gave little weight. Those are separate from her screen-derived material; a refusal to produce a self-report does not mean no inspectable evidence exists. [1]

Keep an observed failure distinct from an unsupported claim or unknown access. Source-support checks also matter when checking stale public records, but they need not reconstruct every earlier action. If the passage can be verified now, how much more evidence does the brief actually require?

A correct substitute can be enough for the brief

The strongest objection is practical: a researcher who can verify the answer should not have to reconstruct the agent's whole history. OpenAI's documented checks concern supporting sources and their dates and authority. [3] I agree that those checks can be enough for a background finding. A verified substitute covering the relevant passage, jurisdiction and period may satisfy the commission, even when the agent's original access remains unknown.

If the brief asked for proposals only, disclosed proposals may also meet its scope. The workbook's admission that it did not observe current portal values is not, by itself, proof of failed proposal work. [1] Requiring an update baseline for every factual summary would add work that the reader's question does not need.

An update can likewise be supported without a complete historical trace. Inspect the relevant baseline and the proposed evidence, then establish the actual difference or correction. If they agree, report no change. The comparison matters because of the claim being made, not because every agent action must be replayed.

The causal objection is equally fair. US/English and Iran/Farsi changed together with the sources and technical environments in Pakzad's experiment. [1] This article does not independently replicate it, and the case cannot tell us which current model is best at multilingual research. It does not show that Farsi caused the weaker result.

Nor would removing an access barrier settle every evidence problem. In Faux Polyglot, Nikhil Sharma, Kenton Murray and Ziang Xiao report preferences for query-language information during retrieval and generation, and for higher-resource-language information during generation when query-language information was absent. Their experiment supplied synthetic multilingual documents. It used English, Hindi, German, Arabic and Chinese, not Farsi, and evaluated a bounded retrieval-augmented generation architecture based on cosine-similarity retrieval. [4]

That study is not a replication of the procurement task. It supplies a separate reason to inspect which available passages support a synthesis: reachability alone does not establish balanced selection. Switching the prompt to English is not a demonstrated fix in these sources. The acceptance decision should follow the evidence needed for the particular brief.

Keep the candidate; leave the update unresolved

Return to D2 and E2. The inspected artifact offers a proposed country name beside an unobserved current portal value, and I2 includes the row for completeness. [1] This article has not verified E2's underlying source. The proposal stays a proposal; it is not an example of a successful correction.

Conditional diagram: unchecked proposal is a candidate; checked proposal can support a fact; an update needs baseline comparison.

Editorial decision guide: verifying a proposal can support a fact; an update claim also needs the relevant baseline and comparison.

The diagram is an editorial decision guide, not measured results. An unchecked proposal remains a candidate. If its supporting passage is verified, it can become a supported fact while its update relationship stays unresolved. Checking both the proposal and the relevant baseline supports an update only if the comparison actually establishes a correction. A checked pair might show no change.

For this workbook, I would preserve the proposal and its source lead. If the commission needs a fact, verify the passage supporting that proposal. If it needs an update, obtain the relevant portal baseline as well. No successful source visit or corrected entry is implied by that recommendation.

An evidence note can preserve the actual finding: the workbook marks the current portal value unobserved and supplies E2 as a proposal. This article has not established that proposal's support or an update. That is suggested editorial wording, not a tested improvement in research outcomes.

The related discussion of retaining a research decision concerns keeping an accepted judgement inspectable later. The discussion of apparent completeness and acceptance evidence addresses the adjacent problem of deciding what a finished-looking deliverable proves. Here, the workbook has already left the missing comparison visible. Keep it visible in the verdict too.

Share this case with a research partner and ask which baseline your next update would need.

References

  1. Roya Pakzad, "Three AI Agents, Two Countries, and One Very Uneven World Wide Web," displayed 3 October 2026. https://royapakzad.substack.com/p/multilingual-ai-agents . Primary account of the procurement experiment, not independently replicated here. Linked artifacts inspected during research on 4 October 2026: https://github.com/royapakzad/llm_agent_world_bank_experiment/blob/main/output/Iran_Claude.xlsx and https://github.com/royapakzad/llm_agent_world_bank_experiment/blob/main/trajectory/claude-iran-recording-transcript.txt . The captured workbook page showed short commit 9b7e552; the transcript page showed 3456b62. Public main links may change. Workbook contents and the screen-derived transcript are inspected artifacts of the same experiment, not independent replications. Proposed factual values and original video were not independently verified for this article.
  2. World Bank, "Global Public Procurement Database: Data Collection Report," historical document with 2020 page headers. https://documents1.worldbank.org/curated/en/384951635848552076/txt/Global-Public-Procurement-Database-GPPD-Data-Collection-Report.txt . Inspected during research on 4 October 2026. The workbook names it as the field-inventory source. It is a foundational collection-method document, not an observation of the live portal baseline or verification of the essay's API chronology.
  3. OpenAI Help Center, "Searching the web with ChatGPT," Review sources and results. https://help.openai.com/en/articles/9237897-searching-the-web-with-chatgpt . Guidance inspected during research on 4 October 2026. Cited for source-support, date and authority checks, and the statement that web-search responses may include citations. Provider documentation, not an independent accuracy measurement or a complete source-access record.
  4. Nikhil Sharma, Kenton Murray and Ziang Xiao, "Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models," NAACL 2025; inspected arXiv v3 dated 11 February 2025. https://arxiv.org/html/2407.05502v3 . Inspected during research on 4 October 2026. Reports source-selection preferences in a bounded synthetic-document setting. Sections 2.1 and 7 delimit the languages and architecture; neither live Iranian portal access nor Farsi was tested.