Skip to content

feat(app): serve a root-domain robots.txt - #551

Merged
Ehesp merged 1 commit into
mainfrom
feat/root-robots-txt
Sep 21, 2026
Merged

Ehesp merged 1 commit into
mainfrom
feat/root-robots-txt

Conversation

@claude

@claude claude Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Requested via Slack thread

One of five PRs replacing #530, which bundled all five root-domain discovery fixes into a single change.

Summary

Before: https://docs.page/robots.txt returns a 404. Per-repo robots files exist only under each hosted repository's own path, so a search or AI crawler arriving at the root domain is told nothing about what it may crawl and is given no pointer to a sitemap.

After: that same URL returns a plain-text policy that allows every crawler and points at the root sitemap:

User-agent: *
Allow: /

Sitemap: https://docs.page/sitemap.xml

In short: the root domain gets an explicit crawl policy and a sitemap pointer instead of a 404. All crawlers, AI crawlers included, are deliberately allowed — there are no AI-specific blocks.

How: a new static App Router route at app/src/app/robots.txt/route.ts holds the policy as a build-time constant, mirroring the existing root llms.txt route. Its cache policy comes from a new ROOT_ROBOTS_TXT_CACHE_HEADERS constant in app/src/proxy.ts (day-long edge TTL, hourly browser revalidation).

Scope

  • app/ (hosted site, MCP, Ask AI)
  • packages/cli/
  • packages/mdx-bundler/
  • docs/ (product documentation)
  • Repo / CI / other

Type of change

  • Bug fix
  • New feature
  • Documentation
  • Refactor / chore

Test plan

  • biome ci . clean (the check CI runs)
  • bun test — 140 pass, 0 fail
  • tsc --noEmit in app/ — clean

Notes for reviewers

🤖 Generated with Claude Code

https://claude.ai/code/session_01Y5HaatshAzKUdXWYUbAC4a


Generated by Claude Code

https://docs.page/robots.txt returned 404. Per-repo robots files exist only
at /{owner}/{repo}/robots.txt, so search and AI crawlers hitting the root
domain got no crawl policy and no sitemap pointer.

Adds a static App Router route serving an explicit "Allow: /" policy plus a
Sitemap directive, with cache policy via a new ROOT_ROBOTS_TXT_CACHE_HEADERS
constant. All crawlers, AI crawlers included, are deliberately allowed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y5HaatshAzKUdXWYUbAC4a
@railway-app

railway-app Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the docs.page-pr-551 environment in docs.page

Service Status Web Updated
docs.page ✅ Success (View Logs) Web Sep 8, 2026 at 10:29 am UTC

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@Ehesp
Ehesp merged commit 9b7bc00 into main Sep 21, 2026
3 checks passed
@Ehesp
Ehesp deleted the feat/root-robots-txt branch September 21, 2026 18:58
claude Bot pushed a commit that referenced this pull request Sep 21, 2026
Brings the branch up to date with main after #550, #551, #552 and #553
landed. `app/src/proxy.ts` auto-merged: the root robots.txt and root
sitemap.xml cache-header constants from main sit alongside this branch's
ROOT_MCP_SERVER_CARD_CACHE_HEADERS, each still consumed by its own route.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y5HaatshAzKUdXWYUbAC4a

This branch was successfully deployed

No deployments
docs.page / docs.page-pr-551 — 9b6635a3 Deployed Sep 8, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants