XML Sitemap and RobotsOptimization
Two small text files control how crawlers see your site, and one wrong line can hide it entirely. We set them up deliberately and test every rule.
XML Sitemap and Robots Optimization: what the work involves
Sitemaps and robots.txt look simple, which is why they cause so many incidents. A staging disallow-all rule gets deployed to production. A sitemap lists redirected, noindexed and error URLs, teaching search engines to distrust it. A plugin generates thousands of useless tag and attachment URLs. A robots rule blocks the CSS and scripts needed to render pages. None of these produce an error message; the site simply performs worse, and nobody connects it to a two-line text file.
We audit both files against the actual state of the site. Sitemaps should contain only canonical, indexable, successful URLs, split into logical files with accurate last-modified dates, and referenced from robots.txt and Search Console. Robots rules should block only what genuinely should not be crawled, tested against real URLs before release. For Next.js and other framework sites we generate both dynamically from your route and content data, so they stay correct as pages are added or removed without manual upkeep.
Core features
Sitemap content audit
Every sitemap URL tested for status code, canonical and indexability, with offenders removed so the file reflects only pages you want kept.
Structure and segmentation
Sitemap index files split by content type, which makes Search Console reports easier to read and indexing problems easier to isolate.
Accurate modification dates
Last-modified values taken from genuine content changes, since inflated dates teach search engines to ignore the field.
Robots rule review
Line-by-line check of directives, wildcards and user-agent groups, with tests of important URLs to confirm they are not blocked unintentionally.
Dynamic generation
Sitemap and robots output built from your routes and content in Next.js or similar frameworks, so changes propagate automatically.
Specialised sitemaps
Image, video and news sitemaps where relevant, plus submission to Search Console and Bing Webmaster Tools and IndexNow notification for Bing.
What we get right before launch
Robots.txt is not a privacy tool
The file is public and blocked URLs can still appear in results if linked. Sensitive content needs authentication or noindex through proper means, not a disallow rule.
Blocked pages cannot be noindexed
If robots.txt prevents crawling, a noindex tag on the page is never seen. Choosing between blocking and de-indexing depends on whether you want crawling stopped or the page removed.
Sitemap inclusion is a hint
Listing a URL does not guarantee indexing. Sitemaps help discovery, but the page must still be useful and technically sound to be kept.
Tools and technology
- Google Search Console
- Bing Webmaster Tools
- IndexNow
- Screaming Frog SEO Spider
- Sitebulb
- Chrome DevTools
- Python
- Schema.org validator
Common questions, answered
What should be included in an XML sitemap?
Only canonical URLs that return a successful status and that you want indexed. Redirects, noindexed pages, error pages and duplicates should be left out, otherwise the sitemap loses credibility.
Can robots.txt remove a page from Google?
No. It controls crawling, not indexing, and blocked URLs can remain listed without a description. To remove a page, use noindex or a proper removal, leaving it crawlable so the tag is seen.
How often should a sitemap update?
Whenever the page set changes, ideally automatically. A dynamically generated sitemap tied to your content data stays accurate without manual effort, which is how we build them for framework sites.
Do we still need a sitemap for a small site?
It is not strictly required but still useful, especially for new sites and for monitoring indexing by section in Search Console. The cost is low, so we generally recommend one.
More SEO services
All SEO servicesIndexing Issue Fixes
A page that is not indexed cannot rank, and the reason is rarely the one people assume. We diagnose each excluded URL and fix the underlying cause.
Crawl Budget Optimization
Crawl budget matters only on large sites, and there it matters a lot. We use server data to see where crawlers spend their visits and stop the waste.
Technical SEO Services
Great content cannot rank if search engines cannot reach it. We find the crawl, index and rendering faults holding your site back and fix them in the codebase.
Ready to start your XML Sitemap and Robots Optimization project?
Tell us what you need and we will come back with a clear scope, timeline and the questions worth answering before any build starts.
