Sample audit: a small guesthouse near Kanazawa Station
Made on 2026-10-04 from public pages only, without the owner's involvement. The business name is withheld because it did not ask for this audit. Every number below was measured on that day and can be re-run. 診断例(日本語の要約は末尾)。
1. How AI answers show the business
| Question | Named? | Notes |
|---|---|---|
| "recommended guesthouse in Kanazawa near the station for solo travelers" | No | Four competitors named instead. |
| "金沢 ゲストハウス 駅近 おすすめ" | Yes, one line | Sourced from listicles, not from the guesthouse's own site. |
| Brand name + "review price rooms" | Yes | All 10 sources were booking platforms; the official site was not cited; prices quoted were platform prices. |
Engine for this sample: Claude with web search, one run per question. The paid audit runs 10 questions in ChatGPT, Perplexity, Google AI Overviews and Claude, with screenshots.
2. Findings
- AI training crawlers are refused at the server, although robots.txt allows them. GPTBot and ClaudeBot user agents get HTTP 403; a browser, OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot get 200. robots.txt only blocks
/wp-admin/. Effect: engines that browse can still read the site, but the models' background knowledge is built without it and leans on booking platforms. (Tested with the published user-agent strings; the owner should confirm in the access log.) - The declared English version is almost empty. The Japanese home says the English page is
/en/. That URL returns a page with a Japanese title and about 680 characters of text. The real English page lives at another URL that no language tag points to. - No structured data. No JSON-LD on either language home: the walking time from the station, dorm and private-room prices, capacity and opening year are written as text only.
- No meta description and no H1 on either home page.
- The official site is not the source for its own facts. In the brand query, answers pointed only to booking platforms, which take a commission on every booking they send.
3. Fixes, in order
- Allow GPTBot and ClaudeBot in the host or security-plugin rule (or keep blocking them as a deliberate choice). About 15 minutes.
- Point the English language tag at the real English page, and add the reverse tag. About 30 minutes.
- Add one
HostelJSON-LD block per language: address, geo, phone, price range, check-in/out, amenities, links to booking-platform profiles. About 30 minutes. - Add a short bilingual FAQ with what travellers ask AI: distance from the station, luggage storage, female dorm, late check-in, whole-house rental. About 1 hour.
- Add a meta description and an H1 to each home page. About 10 minutes.
4. Re-measure
GPTBot and ClaudeBot return 200; the declared English URL carries the English text; JSON-LD present on both homes; the three questions re-run in four engines, counting whether the official site is cited.
日本語の要約
金沢駅近くの小さなゲストハウスの例です。(1) robots.txtでは許可しているのに、サーバーがGPTBot・ClaudeBotを403で拒否。(2) 英語版として宣言している /en/ がほぼ空で、実際の英語ページには言語タグが向いていない。(3) 構造化データなし。(4) 説明文・H1なし。(5) 店名で聞くとAIは予約サイトだけを出典にし、公式サイトは引用されない。直す順番と所要時間は上の「3. Fixes」の通りで、合計2〜3時間程度の作業です。