ANALYSIS · VERIFICATION 16 min read

A public report written with AI: nine errors you can see in the file itself

The model did not fail here. What failed was the step where somebody opens three footnotes and compares the numbers. I pulled the report out of the archive and went through it page by page. Below are nine classes of error, each with a quote. Next to each one, the check that stops it.

The blurred cover of the Polish report „Polaków portret własny” on the right, next to three of its errors set against what its own cited source says

A Polish nation branding report called “Polaków portret własny” (A Self-Portrait of Poles) picked up the label “written by AI” this week, and the discussion stopped there. That is not enough to learn anything from. I opened the file and checked what exactly is wrong with it. Nine repeatable classes of error came out, each with a cheap check, and not one of those checks requires any knowledge of language models.

This is about a document, not a person. Names appear where they belong to the public record of the case. Without them you cannot describe who was responsible for what. I am judging the process that let this file reach the internet.

What actually happened

The document is dated “Warsaw, November 2025” and sat on the website of the Marka Polska Think Tank foundation. Page two carries the grant formula: a project delivered under the public task “Brand Poland. Tourism sub-strategy”, co-financed by the Ministry of Sport and Tourism under contract no. 2025/0014/1251/UDOT/DT/BP/IS of 5 June 2025. The ministry signed that contract not with the foundation but with an association, Konferencje i Kongresy w Polsce. What it bought was research, a strategy and a conference.

The figure in the headlines needs a paragraph of its own, because it is itself an example of what this post is about. The number repeated in the coverage is 200,000 zloty, roughly 54,000 US dollars. It entered circulation through a question Sebastian Meitz put on X: did the ministry pay 200,000 zloty for a report generated by AI. The ministry gave a figure of its own and said it had accepted a settlement of 179,738.19 zloty, roughly 49,000 dollars. The competition the grant came from allowed between 150,000 and 250,000 per sub-strategy. None of those numbers is the price of this file, because the contract covered the whole task. Working out where the headline number came from takes fifteen minutes and changes the sentence.

The report's author is Barbara Mróz-Gorgoń, a professor of marketing, president of the foundation and, since late July 2026, the government plenipotentiary for promoting Poland's brand. On Polsat News she said the report “was certainly assisted by artificial intelligence”. She called it a transitional document and a work in progress that should never have reached the public domain, and added that she takes the blame. In a written statement on X she said the role of generative AI was “deliberately limited to an auxiliary, non-authorial stage”. Full responsibility for the expertise, the editing and the substantive verification, she wrote, rests with the authors.

The ministry distanced itself on three levels. It did not order the report and does not count it as the delivery of any contract. No request came in to use its logo. The statement that the report was financed from ministry funds is, according to the ministry, untrue. That statement is printed on page two of the report. The story broke on 20 August 2026 with a post by Sebastian Meitz of the Sobieski Institute. The file disappeared from the foundation's site within a day, and the address now returns a 404. Łukasz Olejnik pointed out that the document “often moves from what one person said to a claim about entire nations, and then to a psychological diagnosis”. He also listed factual errors, contradictions in the description of its own study, and numbers with no sources. The rest of this post checks those charges against the file and turns them into a procedure.

Nine errors you can see in the file itself

Every quote below comes from the PDF, 26 pages, in the version archived on 14 May 2026. Original spelling, capitals included, with a translation where the Polish carries the point.

1

A number that is not in the source it cites

“GDP per capita grew from about 12.5 thousand USD in 2010 to about 17–18 thousand USD in 2024”

The footnote points at the World Bank, indicator NY.GDP.PCAP.CD. That same indicator gives 25,104 USD for Poland in 2024. The 2010 figure matches exactly. This is the hardest variant to catch, because the source is real and checkable. Only the number is wrong.

2

A number with no footnote at all

“The young generation of Poles speaks FLUENT ENGLISH (95%), often ADDITIONALLY GERMAN, FRENCH or MANDARIN”

That number has no reference and no basis. Same for “60% of Americans cannot point to Poland on a map”, the only bullet on its list without a footnote marker. In a document like this, a striking number with no source is the rule rather than the exception.

3

A footnote that names an institution instead of a document

“UNWTO and OECD – reports on the growth of nature tourism, slow travel and culinary tourism”

That exact footnote sits under four different numbers and four different claims. The bibliography runs to a dozen or so entries and almost every one says “(n.d.)” instead of a year, with no title of any specific publication and no page. Two of them break off at a city name: “Copenhagen:” and “Washington, DC:”. It looks like sourcing and is not sourcing.

4

The same fact in three versions in one file

cover: “7 Design Thinking workshops in 7 different cities” · methodology: five cities, named · summary: “Design Thinking workshops in 3 cities”

The sample has the same problem. “35 FGI” on the cover becomes “35 people in total” in the methodology section, so participants rather than focus groups. The breakdown of 24 interviewees adds up to 25. Nobody read this document end to end in one sitting.

5

A jump from one remark to an entire nation

“One respondent recalled a conversation with the therapist of a friend from Germany […] This constitutes a fundamental DIFFERENCE in models of social bonding between Poland and Western Europe”

A second-hand anecdote about one person is promoted to a cultural difference between Poland and the rest of the continent. A model will happily close a sentence like that, because closing sentences is its job. Deciding how much weight a remark carries is the researcher's job.

6

A clinical diagnosis standing in for a finding

“IMPOSTOR SYNDROME at a NATIONAL LEVEL” · “STOCKHOLM SYNDROME – THE PARADOXICAL LOVE OF ONE'S HOMELAND”

Clinical terms describe a whole country here, with no instrument, no sample and no result. They read like findings and they are metaphors. In a report paid for with public money, that is a sentence you either cut or back with a study.

7

A source that has not existed for four years

“In the World Bank's «Doing Business» ranking, Poland ALWAYS ranks WORSE than its level of economic development would suggest”

The World Bank shut Doing Business down in September 2021 after data manipulation was found. The model writes about it in the present tense, because in its training data the ranking is alive. “Always ranks worse” is also phrased so that nobody can check it.

8

Geography nobody checked against a map

“Olsztyn (north east), Szczecin (west), Katowice (northern centre) and Wrocław (north west)”

Katowice is in the south, Wrocław in the south west. The press caught the first error; the second one sits in the same sentence. Checking both takes ten seconds and needs no source beyond a map.

9

Generation artefacts nobody removed

“Despite RZECZYWISTYCH OSIĄGNIĘĆ” · “w historii sufferingu” · “the archetype of the CARER and the GUARDIANKA” · “Polskie płaskiz”

An English word left inside a Polish sentence, an English noun given Polish grammar, an invented word, and a word cut off mid-spelling on a list of Poland's natural assets. Footnote numbers collide into strings like 1611, 2114 and 2316. The same respondent quote appears twice in two different wordings, though a quote has one wording by definition.

WORK WITH ME

This is what I do hands-on: advising on AI strategy and building agents that survive the demo.

The model had no way to turn a light red

None of those nine points is a tool failure. The model did exactly what it is for: it produced text that looks like a report. A report has footnotes, so it got footnotes. A report has numbers, so it got numbers. The shape is right and the backing is not, and “I do not know, I do not have that number” is not a shape a report takes.

That asymmetry is why documents get hit hardest. A developer whose model invents a function gets a red test and a stopped pipeline. A person writing prose gets a smooth paragraph and nothing else. There is no signal, because in prose there is nothing to run. I described the same mechanism when writing about the confidence figure a model attaches to its own answer. It is a tone of voice, not a measurement of correctness. It rises with the fluency of the sentence rather than its truth.

So the only filter left is what the author already knows about the subject. With GDP per capita even that is not enough, because almost nobody carries the value of that indicator in their head. What is enough is an action: open the source and compare the number.

A draft is not a defence

A draft is a file that sits with its author. This report carries a logo and a contract reference on page two. It went out under the banner of a publicly funded project. It stopped being a draft the moment somebody clicked publish. A label applied afterwards does not undo that. Only a correction does, together with a note saying exactly what was fixed. There is also the matter of time. The file path on the foundation's server points to November 2025. The Internet Archive has captures of the report page from 10 December 2025 on. The file came down in the second half of August 2026. The draft sat in public for over eight months.

The question of responsibility is more interesting than the question of the tool. Three parties stand next to each other here: the ministry that gave the grant, the association that signed the contract, and the foundation that published the report. The last two links are not as far apart as they look from outside, because the report's author is also the foundation's president. The ministry says it did not order this document. The document says it was made with ministry funds. The author's statement says substantive verification sat with the authors. Three sentences in which nobody owns the publication. The file went out anyway. Who signs a document before it is sent is the first question in the register of AI uses a company has to keep. It is also the last one anybody asks before publishing.

Polish public administration has this written down, as it happens. The Ministry of Digital Affairs published a guide to artificial intelligence for public administration in March 2026. It puts the rule plainly: final responsibility for the content of a letter, a communication or an administrative decision always rests with a human being, not with an AI system. It also asks staff to check every legal act or document the system cites by opening that document themselves. The guide came out a few months after this report. The rule it writes down is not new.

This is not a Polish speciality

In October 2025 Deloitte Australia returned part of its fee for a 237-page report written for the country's employment department. The contract was worth 440,000 Australian dollars and the refund covered the final instalment, over 97,000. The report carried references to publications that do not exist. Among them were invented papers by a researcher who does exist, and a fabricated quote from a federal court judgment. An outside academic caught it, not the client and not the supplier. He opened the citations. The corrected version now carries a note that a generative language model was used in the writing.

In the United States the same mechanism reached the courts. In June 2023, in Mata v. Avianca, a judge fined two lawyers and their firm 5,000 dollars. Their brief stood on six decisions that never existed, complete with quotes and docket numbers. Damien Charlotin at HEC Paris maintains a database of such rulings. On the day I am writing this, it holds 1,936 cases from around the world.

The common denominator is not the industry and not the country. It is a document that travelled from prompt to reader with no step for comparing a claim against a source. I see the same pattern in code. a test suite can run green while asserting nothing. A footnote can look like a source the same way, while being none. In both cases you are not measuring what you think you are.

The verification step, as a one-page procedure

Verification as a slogan changes nothing. Nobody knows what to do on Monday morning. Below are nine actions, one for each class of error above. A person does all of them and none needs a tool. On a twenty-page document they take under an hour together.

1. Open three footnotes at random

Do not check that the source exists. Check that the number is in it. This is the only control that catches error number one.

2. Every number has a footnote or it goes

No exemption for “commonly known” figures. A percentage with no reference gets deleted or rewritten as a sentence without a number.

3. A footnote names a document, a year and a page

The name of an institution is not a source. If the footnote does not lead to one file, there is no footnote.

4. Search your own file for every number

Ctrl+F through your own document. The summary, the body and the conclusions have to carry the same value in the same words.

5. Count the people behind each generalisation

How many said it. If one, the text has to say “one interviewee”, not “Poles”.

6. A clinical term needs an instrument

Syndrome, disorder, trauma. Either a measurement stands behind it, or it goes back to the metaphor drawer and out of the findings.

7. Check the date of a source, not just the name

Rankings, statutes and annual reports expire. The model will not notice, because in its data they are still running.

8. Ten-second facts get ten seconds

A map, a date, the spelling of a name, a count of laureates. These cost the most credibility for the lowest cost of checking.

9. Somebody else does the last read

A person who did not write the text reads the whole thing in one pass. Only they will see an English word sitting inside a Polish sentence.

Two process rules go with that. First: the document carries, in its own text, the name of the person who checked it and the date. Without that, verification is a declaration rather than an event. Second: do not ask a model to check its own text. A model grading its own work approves it almost every time. It judges the same sentence with the same probability distribution that produced it.

How I drill this in a training room

This procedure emailed to a team changes nothing. People start using it once they watch the tool get something wrong on their own material. So a training day for a department that publishes text is built around two rules.

First: we work only on material the participants know by heart. Their report, their request for proposal, their job description. On somebody else's example nobody spots the error, because there is nothing to compare it against. The same rule holds on the technical side, where the workshop runs on the team's own repository rather than a sample project. There a red test does part of the work for the trainer.

Second: I break the tool on purpose. I do not wait for it to break by itself on a real publication. I ask for a question from the participants' own field, one the model ought to fail. It usually does. Then we take the answer apart in front of the room, in three columns. What is true. What merely sounds true. How you tell one from the other without leaving your desk. That last column stays with people longer than any lecture about models.

Then comes a forty-minute exercise lifted straight out of this report. Everyone takes a document of their own from a year ago. They ask the model for a summary with figures. Then they open three of those figures in the original and compare. So far, in every room, somebody has found error number one in their own work. What such a day looks like from the organisational side I wrote up after a session for the newsroom at the TVP Television Academy. That post also lists the failures that can wreck it.

The section above describes a room that publishes text. A team that publishes code has the same problem in a different place and needs a longer format. The programme of the two-day workshop, the terms and the enquiry form live on the training page for technical teams.

The test you can run before buying anything is one sentence long. Ask whose material the participants work on. If it is a sample, you are buying a demonstration.

Frequently asked questions

Can you tell that a text was written by AI?

AI detectors produce too many false hits to base a decision on. The only reliable signals are material ones. A number that disagrees with the source it cites. A footnote naming an institution instead of a document. The same claim in two versions in one file. An untranslated word inside a sentence. You check the content, not the style. Style can be fixed in a minute; a number that disagrees with its source stays.

Is it allowed to use AI for official documents and reports?

It is, with human oversight. The Polish Ministry of Digital Affairs guide for public administration says outright that responsibility for the content of a letter always rests with a human. Every source the system cites should be opened by hand. The problem in this case was not the use of a tool. It was the absence of the step where somebody does that.

How long does verifying a twenty-page report take?

Under an hour for a document with a dozen or so footnotes. One condition: somebody who knows the subject does it, working from a list of actions rather than a general instruction to “check this”. The expensive part is action number one, opening the sources and comparing figures. The rest is reading with a pencil.

Can another model verify the text instead of a person?

It can prepare the list of claims to be checked, and that is its best use here. It cannot confirm that a number matches a source it never opened. A model asked to assess its own text approves it in the large majority of cases. The second model's role ends at preparing work for a human.

SP

Szymon Paluch

Claude Certified Architect · ex-CTO

Does your writing go out without that step?

Tell me which document your team publishes most often. Fifteen minutes is usually enough to settle whether there is a training day to build from it, and where the check belongs.

Book a call
Related posts
What an AI agent costs and what actually drives the quote
Claude Certified Architect certification: the complete guide
I passed Claude Certified Architect: decisions, not definitions