
Making your website readable by AI: robots.txt, JSON-LD and llms.txt
When a recruiter asks an AI assistant "Who is Mohand Oussadi?" or "Which Next.js developers have multi-tenant experience?", the answer depends on what AI crawlers could read on the web. SEO is no longer just about Google: your site also needs to be readable by AI.
While auditing this portfolio, I found several problems that made part of my content invisible. Here is the approach, layer by layer.
Layer 1: allow access (robots.txt)
AI crawlers generally respect robots.txt. Each provider has its own agents: GPTBot and OAI-SearchBot (OpenAI), ClaudeBot and Claude-SearchBot (Anthropic), PerplexityBot, Google-Extended… Allowing them explicitly removes any ambiguity.
My first finding was brutal: on the live site, /robots.txt returned a 500 error. The dynamic [lang] route caught the URL and crashed, since "robots.txt" is not a language. And when robots.txt returns a server error, Google temporarily treats the whole site as disallowed.
// app/robots.ts
import type { MetadataRoute } from "next";
export default function robots(): MetadataRoute.Robots {
return {
rules: [
{ userAgent: "*", allow: "/" },
{ userAgent: ["GPTBot", "ClaudeBot", "PerplexityBot"], allow: "/" },
],
sitemap: "https://mohand.oussadi.com/sitemap.xml",
};
}The full fix: this file, plus export const dynamicParams = false in the [lang] layout so any unknown language returns a clean 404 instead of an error.
Layer 2: list the pages (sitemap.xml)
A sitemap generated from the same data as the site (pages, articles, French and English versions with their alternates) guarantees no page is forgotten and stays up to date effortlessly.
Layer 3: content in the HTML
This is the most underestimated point. Most AI crawlers do not execute JavaScript: they read the HTML the server returns, full stop. Anything not in it does not exist for them.
Next.js server rendering handles most of it… except hidden content. My tabs and accordions (Radix UI) did not render closed content: the details of my experience and my entire education were missing from the HTML. The fix: forceMount, which keeps the content in the HTML while hiding it on screen.
<AccordionContent forceMount className="data-[state=closed]:hidden">
{job.responsibilities.map((r) => <li key={r}>{r}</li>)}
</AccordionContent>To check, a simple curl https://your-site.com | grep "expected text" is enough: if the text does not show up, crawlers do not see it.
Layer 4: structured data (JSON-LD)
HTML says what to display; structured data says what it is. A JSON-LD block using the schema.org vocabulary unambiguously describes a person, their job, skills and education.
{
"@context": "https://schema.org",
"@type": "Person",
"name": "Mohand Oussadi",
"jobTitle": "Full-Stack Developer",
"sameAs": ["https://www.linkedin.com/in/mohand-oussadi-091777106/"],
"knowsAbout": ["React", "Next.js", "Node.js", "TypeScript"],
"alumniOf": { "@type": "EducationalOrganization", "name": "Université Paris 12" }
}Generate it from your existing data rather than by hand: it will never drift from the visible content.
Layer 5: a summary for AI (llms.txt)
Proposed in 2024 by Jeremy Howard, llms.txt is a Markdown file at the site root that summarizes its content for language models: who, what, key links. It is not yet a universally used standard, but it costs almost nothing. On this site, it is generated at build time from skills, experience, projects and articles.
And the rest of SEO hygiene
- Unique titles and descriptions per page.
- Canonical and `hreflang` tags for a multilingual site.
- No keyword stuffing: the
keywordstag is ignored by search engines, and a 70-term list looks like spam.
Conclusion
Being readable by AI rests on the same foundations as good SEO, with one extra requirement: everything must be in the HTML. Check that robots.txt responds, that your content shows up in a plain curl, describe yourself with JSON-LD and add an llms.txt. A few hours of work so assistants talk about you accurately.