Call · 15 min
StudyMattia Esposito26 September 202612 minute read

The website lets everyone read it. One producer in six offers the technical sheet.

We read from outside the 249 websites declared by producers and exporters in the ICE showcase for foreign buyers. The websites let everyone read them, search engines and AI assistants included. The technical sheet a buyer wants to see is offered by one food producer in six.

In short

19 food websites out of 121, 15.7%, offer a technical sheet to download between the home page and a product page, and another 7 show it only on the page. Among wineries 13 out of 25 offer it, among other food producers 6 out of 96. Every sheet was opened and read, one by one.

121 food websites out of 121 let in Google, Bing and the search crawlers of ChatGPT, Claude and Perplexity. The gap is on the page: in 18.2% the title is only the company's name, 35.5% have no meta description, 43.8% do not declare with structured data who the company is, 11.6% have an email address invisible to AI assistants.

26 addresses out of 249 in the showcase, 10.4%, do not lead to a working website: domains that no longer exist, error pages, sites under maintenance, a WordPress never set up, a domain that passed to an online casino.

The study is by Itria AI (itria.io), independent of ICE: the directory is the public source the website addresses were taken from. The measurements are of 26 September 2026. Method, checks, corrections and the data for each site, without company names, are on this page and in the downloadable file at the bottom, under a CC BY 4.0 licence. For exporters, the operational pages start from the guide to operational export for the small food producer.

Where export trade really happens

In person, at trade fairs such as Anuga and SIAL, and with samples. The CBI, the Dutch government centre that helps exporters from middle and low-income countries sell in Europe, writes it in its guide on how to find European buyers of processed fruit and vegetables: most lasting partnerships still arise from personal introductions, sample analysis and factory visits.

«Processed food is a face-to-face business, but a website gives you global visibility» (CBI, guide on European buyers of processed fruit and vegetables, updated on 19 September 2023)

Before the meeting comes the search. European buyers look for new suppliers on Google, with the product name and words like «producer» or «supplier», sometimes together with the country of origin, and for the CBI appearing on the first page for those searches is «very important». When it is the producer who comes forward, it recommends a phone call followed by an email.

On the website, according to the CBI, product documentation should be kept available to buyers: complete specifications, including size, weight, packaging, shelf life and nutritional values, that is the information the buyer wants to see and that helps them decide. As an example it cites a producer that publishes a PDF sheet for every product: it is the technical sheet this study looked for.

The study measures what can be seen from outside: what a buyer, a search engine and an AI assistant find on the website. How many export requests arrive from the website is known only to the company that receives them. In the public sources we read that number is missing, and on this page we do not estimate it.

What we measured: the buyer, the search engine, the AI assistant

A company website today has three readers. The foreign buyer who opens it, the search engine that indexes it, the AI assistant that reads it on someone's behalf. We measured five things the first one needs and six that decide whether the other two find the company and understand what it does.

MeasurementWhat counts as yesWhy it matters
English version

English hreflang, link to /en/, EN selector, or home page already in English.

It is the foreign buyer's language.

Contact channel

A form with an email or message field, a mailto, a written address.

Without it, the buyer does not write.

Email readable by an AI assistant

The address in the clear in the served HTML, not obfuscated.

Assistants that read on the fly do not run JavaScript (Vercel).

Technical sheet

To download: a product sheet document, PDF, image or web page, linked between the home page and a product page and opened. On the page: a section titled technical sheet or technical data, with the fields below. A general catalogue does not count.

The CBI recommends keeping it on the website, available to buyers.

Conditions for foreign buyers

An export or wholesale price list, a minimum order for resellers, an Incoterm, an ex works price. Read case by case.

Without them, every offer starts from an email with a question.

robots.txt open to Google and Bing

Googlebot and Bingbot can read the home page.

Whoever blocks them drops out of search results.

robots.txt open to AI search

OAI-SearchBot, PerplexityBot, Claude-SearchBot and Claude-User can read the home page.

OpenAI: whoever blocks OAI-SearchBot does not appear in ChatGPT's search answers.

Indexable home page

No noindex in the meta robots or in the X-Robots-Tag header.

Google removes a page with noindex from results entirely.

Title that says what the company does

The title contains something beyond the company's name and words like Home or Welcome.

Google asks for descriptive titles and is explicit: no «Home».

Meta description

Present and not empty.

It is one of the sources of the text that accompanies the result.

Company structured data

The home page declares schema.org Organization or LocalBusiness.

They help Google understand who the company is and tell it apart in results.

The reason for each row comes from the documentation of those who decide: the pages of OpenAI, Anthropic and Perplexity on their crawlers, Google's guides on noindex, title and organization data. We also measured the sitemap and product structured data, reported separately because none of these sources asks a small site for them.

Why we did not measure llms.txt

Because Google says it does not use it. In its guide to generative search features, updated on 10 July 2026, under the item on llms.txt files: no special files are needed to appear in search, not even in AI features, and creating them «will neither harm nor help» visibility, because Google Search ignores them.

«You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn’t use them.» (Google Search Central, guide to generative search features)

itria.io also has an llms.txt file, meant for other services that read it. We say so because the study's criterion applies to us too: we measure what search engines and assistants declare they use. The same guide adds that structured data is not a requirement for generative search, and that is why we present it for what it does according to Google: say who the company is.

The method: the ICE showcase, the served HTML, the product page in the browser

The frame is ICE's public directory «Find your Italian partner», designed for foreign companies and agents looking for Italian products or partners. The listings were collected between 19 and 31 August 2026 with searches for words linked to Puglia, such as Puglia, Apulia, Bari and Salento, and by product, such as preserves and cheese.

The directory does not record the region and the search matches the text of the listings, so the companies found come from all over Italy. The listings declared 249 different websites, and we read them all, without sampling. The sample represents those who present themselves to foreign buyers in that showcase, and Italian food exports as a whole remain outside its reach.

For each site we read robots.txt, the home page, the contact page and the products page if linked from the home page, on 26 September 2026 between 13:34 and 14:39. We read the HTML the server delivers, that is what an AI assistant that does not run JavaScript reads, with a tool that declares its identity, respects robots.txt and waits at least 2.5 seconds between one request and the next.

Each measure has three outcomes: yes, no, and not measurable. A site that refuses the reading or answers with an error is counted separately and stays out of the percentages. Pages, headers and robots.txt were saved, so every measure can be rechecked on the code read that day.

For the technical sheet we went one click further. On the evening of 26 September we opened in a browser, one site at a time, the first product page of every food website and up to two pages for professionals or documents linked from the home page or the products page. Every candidate document was opened and read: general catalogue, brochure and price list do not count.

The showcase: 26 addresses out of 249 do not lead to a working website

10.4% of the addresses the ICE showcase shows foreign buyers lead to something that is not a company's website. 14 domains no longer exist, 2 home pages answer with an error, and 10 pages are something else: a server default page, two sites under maintenance, a WordPress never set up, a parked domain, a domain that passed to an online casino.

Outcome on the 249 addressesAddressesShare
Read

195: 121 food websites on 122 addresses, 71 from other sectors, 1 not classifiable, 1 shop under maintenance

78.3%

Almost empty home page in the HTML

14: 5 built in JavaScript, 7 that are not a website, 2 minimal but working

5.6%

Domain that no longer exists

14

5.6%

Home page in error

2

0.8%

robots.txt unreachable or in error

14, counted as not measurable

5.6%

Reading forbidden or refused

6

2.4%

Other not measurable

4: two sites that answer a browser and not the tool, a domain back online, an address misspelled in the listing

1.6%

The 10 pages that are not a website, one by one: a server's default page, a server placeholder, a site under maintenance, an online shop under maintenance behind a password page, one being updated, a WordPress with the sample page and noindex on, a parked domain, a redirect to a domain that no longer exists, a domain that passed to an online casino site, a home page made only of PHP error messages. All addresses without a website were reopened in a browser on 26 September.

Whoever curates a showcase of companies for foreign buyers has here a check that takes an hour: open the addresses. For the company, the website declared in a directory is a promise made to a buyer it does not yet know.

The table: 121 food producer websites

Of the 195 sites read, 121 belong to producers or distributors of food and beverages. We classified them all one by one, reading title, description and home page text: the directory also returns companies from other sectors, from cosmetics to machinery, because the search is text-based. The percentages are calculated on the 121 websites: two addresses in the showcase lead to the same winery.

What the buyer findsWebsites with a yesOut of 121
Contact channel in the code

120

99.2%

English version

93

76.9%

Email readable by an AI assistant

107 (7 obfuscated, 7 missing from the code)

88.4%

Technical sheet to download, between the home page and a product page

19

15.7%

Technical sheet on the page, with the fields

10 (7 without the sheet to download)

8.3%

Conditions for foreign or wholesale buyers

1 (2 counting a borderline case)

0.8%

The foreign buyer almost always finds a way to write. The technical sheet can be downloaded on one website in six, the conditions for ordering on one in 121. Every piece of missing information becomes an email with a question.

And response time matters: in the 2007 InsideSales.com and MIT study, whoever calls back within five minutes is up to a hundred times more likely to reach the contact, as we explain in the five minutes that decide a request.

The cost of the questions that arrive by email adds up to hours nobody invoices: the sum of what replying by hand costs has six lines. Two documents are enough to join those who offer them: the food product technical sheet template, with its English version field by field, and the export price list with EXW and DDP.

Where the technical sheet is: one winery in two, one other producer in sixteen

The downloadable technical sheet is concentrated in wine: 13 wineries out of 25 offer it, 52.0%, and 6 of the other 96 food producers, 6.3%. It is almost always on the single product page, one click after the home page: 17 cases out of 19. In the other 2 it is on the page that lists the wines.

Technical sheetWineries, 25Other food producers, 96
To download

13 (52.0%)

6 (6.3%)

On the page, with the fields

5 (20.0%)

5 (5.2%)

At least one of the two

16 (64.0%)

10 (10.4%)

We made the split between wineries and other food producers after the reading, when we saw where the sheets were, with a rule written before counting: a winery is a site whose main product is wine, and companies that make oil and wine stay among the others. The downloadable file has the column, so the split can be rechecked.

Outside wine, the 6 downloadable sheets belong to three olive oil producers, one truffle producer, a confectionery company and a dairy. Of the 19 documents almost all are PDFs: in 3 cases the sheet is an image or a scan, in 1 it is a wine's electronic label, a web page linked as «Technical sheet».

Who has the sheet and does not show it

On 3 websites the technical sheet exists but cannot be downloaded: a page of sheets protected by a password, a download area that asks for a login, a «Technical sheet» link that opens a form to receive it by email. We did not count them among the yeses, and no form was filled in.

On 4 websites the price list, the prices or the product list arrive only after a form or a registration. On 11 the only product document is a general catalogue or brochure, and on 28 there is no single product page: the products are all on one page or in a gallery, or not listed at all. Of the 27 pages for professionals or documents read, only one offers the sheets to download.

Search engines and AI assistants: robots.txt is open, the page says little

121 food websites out of 121 let in Googlebot, Bingbot and the four crawlers with which ChatGPT, Claude and Perplexity build search answers. None blocks even the training crawlers, and none uses the robots.txt managed by Cloudflare, which according to its documentation closes only training. The door is open: what is missing is inside the page.

What search engines and assistants findWebsitesOut of 121
robots.txt open to Google and Bing

121

100%

robots.txt open to AI search crawlers

121

100%

Indexable home page, without noindex

121

100%

Page served over HTTPS

120

99.2%

Title made only of the company's name

22

18.2%

No meta description

43

35.5%

No structured data on who the company is

53

43.8%

Product structured data on the pages read

9

7.4%

Sitemap declared or present

113

93.4%

A title like «Home - Company name» tells the search engine and the buyer only what the company is called, and Google explicitly asks to avoid «Home». A home page without Organization data leaves Google the job of working out who the company is and telling it apart from a namesake. These are fixes for an afternoon, and they work for all three of the site's readers.

With assistants the game is measured in citations, because the click rarely comes: the Pew Research Center observed that people who see an AI summary on Google click a result in 8% of visits, against 15% of those who do not see it. How to work on both fronts is explained in getting found on Google and ChatGPT.

The email address the AI assistant cannot see: 11.6%

14 food websites out of 121 have no email address readable by an AI assistant: 7 obfuscate it, 7 do not write it in the code. With Cloudflare's obfuscation, for example, the address in the HTML becomes «[email protected]» and returns to the clear only with a script in the browser, as Cloudflare's documentation explains. The itria.io case, and the check to run on your own site, are in the lab piece on the email AI assistants cannot see.

AI assistant crawlers read the HTML as it arrives: Vercel measured it on about 1.3 billion requests in a month, and none of the major AI crawlers runs the pages' JavaScript. Whoever asks an assistant for that company's contact finds the website and does not find the address.

What we corrected before publishing

A first reading, on 25 September, gave 2 technical sheets and 5 pieces of information for foreign buyers out of 128. Checking the positive cases one by one we found two rules that were too broad: one counted the word «Wholesale» in a menu and a retail cart's minimum order, the other classified a vending machine manufacturer as food. We reread all the websites with the corrected rules.

MeasurementFirst reading, 25 SeptemberSecond reading, 26 September
Food websites

128, classified by keywords

123, classified one by one

Downloadable technical sheet

2

1: the other belonged to a vending machine manufacturer

Conditions for foreign buyers

5

1: the others were the word Wholesale in a menu and in a presentation, a spam text, a retail cart

Almost empty home page

14, all attributed to JavaScript

14: 5 in JavaScript, 9 that are not a website

Then a third reading, on the evening of 26 September. The second had looked at three pages per site and counted 1 technical sheet out of 123. Opening a product page in the browser too, the downloadable sheets are 19, and rereading the websites we found two errors in the count of food websites: all percentages are now calculated on the 121.

MeasurementSecond reading, 26 September afternoonThird reading, 26 September evening
Food websites

123

121: two addresses lead to the same winery, and an online shop under maintenance moved among the addresses without a website

Technical sheet to download

1, on three pages per site

19, opening a product page too. Already within the three pages there were 3: the rule did not recognise «factsheet» or a sheet as a web page

Addresses without a working website

25 out of 249

26 out of 249

The check, and how far it can be trusted

Every "yes" for technical sheet and conditions for foreign buyers was read one by one in the saved code, and every measured site was classified one by one. This case-by-case reading was done by an Itria AI agent, on the code saved on the day of the measurement, with rules written beforehand. The positive cases and the addresses without a website were then reopened in a browser: the check confirmed the positive cases and removed 7 addresses from the showcase count, including a page that redirects to the brand's website and two sites that answer a browser. The borderline case is an FAQ for importers that explains what the minimum order depends on without stating it: we counted it as no, and counting it the figure goes from 1 to 2 out of 121.

In the third reading every sheet was opened and read: the PDFs with a program that extracts their text, scans and images looked at one by one, the sheets on the page read in the code of the page opened in the browser, label and fields. On 1 site the fields of the sheet on the page were not in the code: we counted it as no.

robots.txt and noindex were checked on all files, and the tool recognises a block on a test robots.txt. Title: all cases reread one by one. Description, structured data and sitemap: 20 sites drawn at random with a fixed seed, reread in the code, identical outcome in 20 cases out of 20. With 20 cases the true error rate of these three measures can reach about 15%, according to the rule of three.

One limit of scope remains. For the sheet we opened one product per site, the first in the order of the products page: a site that puts the sheet only on some products may come out without one, and a sheet behind restricted access is not counted. The other measures stay on the three pages read on the afternoon of 26 September.

Downloading the data, and how to cite it

The file has one row for each of the 249 addresses, with date, outcome, sector and all the measures, including those of the third reading on the product page, without company names and in shuffled order. The data is licensed under CC BY 4.0: it can be reused freely by citing Itria AI (itria.io) and the link to this page.

FileWhat it containsLink
Study dataCSV, 249 rows, CC BY 4.0

Date, outcome, sector, the measures for the buyer and for search engines and AI assistants, the pages read, and for food websites the product page, the sheet to download, on the page or hidden, and whether it is a winery.

studio-siti-produttori-alimentari-2026.csv

Questions and answers

How many food producer websites offer a downloadable technical sheet?

In Itria AI's study of 26 September 2026, 19 websites out of 121, 15.7%, between the home page and a product page, and another 7 show it only on the page. Among wineries 13 out of 25 offer it, among other food producers 6 out of 96. These are the websites of food producers in the ICE directory for foreign buyers.

Every sheet was opened and read, one by one. The count looks at the first product of each site: a site that puts the sheet only on some products may come out without one.

Do food producers block Google or AI assistants in robots.txt?

No: in Itria AI's study of 26 September 2026, 121 food websites out of 121 let in Googlebot, Bingbot, OAI-SearchBot, PerplexityBot, Claude-SearchBot and Claude-User, and none blocks even the training crawlers.

The gap is on the page: in 18.2% the title is only the company's name, 35.5% have no meta description and 43.8% do not declare with structured data who the company is.

Do you need an llms.txt file to appear in Google's AI answers?

Google says no. In its guide to generative search features, updated on 10 July 2026, it explains that no special files such as llms.txt are needed to appear in search, AI features included, and that creating them neither helps nor harms visibility, because Google Search ignores them.

That is why Itria AI's study measures what search engines and assistants declare they use: robots.txt, noindex, title, meta description and structured data.

Why can't an AI assistant find a company's email address on its website?

Because the email address is obfuscated or not written in the HTML. Assistants that read a page on the fly do not run JavaScript and see only the HTML the server delivers. In Itria AI's study, 14 food websites out of 121, 11.6%, had no readable email address: 7 obfuscated it and 7 did not write it in the code.

How was Itria AI's study of food producer websites carried out?

All 249 websites declared in the ICE Find your Italian partner directory listings, collected with searches for words linked to Puglia and by product, were read. For each site robots.txt, home, contact and products pages were read in the served HTML, on 26 September 2026, respecting robots.txt and without filling in forms. For the technical sheet, the first product page was also opened in a browser.

Every site was classified one by one and every positive case of technical sheet and conditions for foreign buyers was read case by case. The data, without names, can be downloaded under CC BY 4.0.

Notes on sources

  1. ICE, Find your Italian partner: the public directory for foreign companies and agents looking for Italian products or partners. It is the study's frame: the data on the websites are Itria's measurements, and the study is independent of ICE.
  2. CBI, the centre for the promotion of sustainable production and trade between middle and low-income countries and Europe, part of the Netherlands Enterprise Agency and funded by the Netherlands Ministry of Foreign Affairs: 7 tips for finding buyers in the European processed fruit and vegetables market and 7 tips for doing business with European buyers of processed fruit and vegetables, both updated on 19 September 2023. The English sentences are quoted verbatim.
  3. OpenAI, OpenAI's crawlers: «Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers». For ChatGPT-User robots.txt may not apply, which is why it is not among the measures.
  4. Anthropic, Anthropic's crawlers, updated on 7 April 2026: disabling Claude-SearchBot or Claude-User may reduce the site's visibility in users' searches; ClaudeBot concerns training.
  5. Perplexity, Perplexity's crawlers: to appear in results it recommends allowing PerplexityBot in robots.txt.
  6. Cloudflare, managed robots.txt, updated on 3 August 2026: it blocks training crawlers such as GPTBot, ClaudeBot and Google-Extended; and email address obfuscation.
  7. Vercel, The rise of the AI crawler: the major AI crawlers do not run JavaScript.
  8. Google Search Central: noindex, title, organization structured data (updated on 8 September 2026), sitemap and guide to generative search features (updated on 10 July 2026). The quotation on llms.txt is taken verbatim.
  9. Pew Research Center, clicks with and without an AI summary, 22 July 2025: 900 US adults, 68,879 searches in March 2025, a click on a result in 8% of visits with a summary and in 15% without.
  10. The measurements are by Itria AI (itria.io), taken on 26 September 2026 between 13:34 and 14:39 with a tool that reads the served HTML, respects robots.txt and declares its identity; the product page was read in a browser on the evening and night of the same day. All the sources above were opened and read on 26 September 2026.
  11. No company name is published, either on the page or in the file. No request was sent to the companies measured.
·The next step

Seeing your own website the way its three readers read it.

The same measurements can be taken on your website: what the buyer finds, what Google reads, what an AI assistant understands. Write us a line about what weighs on you: we take the first step, even if we end up not working together.