Technical SEO
Why our sitemaps come from the codebase
A sitemap should describe the website that actually exists. That sounds obvious, but hand-maintained XML files drift surprisingly fast.
A new page gets launched and never makes it into the sitemap. An old route is deleted but remains listed. A rebuild changes the URL structure while the XML file continues describing the previous site. Every manual list becomes another place where reality and documentation can diverge.
We prefer to generate sitemaps from the same sources that generate the site.
The route system already knows the inventory
On one of our larger service sites, the sitemap is assembled from several sources:
- a small explicit list of true static pages
- the product or service definitions used by the site
- the location content loader
- the article content loader
The sitemap does not need a second manually maintained catalog of those pages because the application already has one.
If a valid content file creates a location page, the sitemap can read the same collection. If an article disappears from the content directory, it disappears from the sitemap too.
That reduces drift.
Static pages still deserve explicit treatment
Not every route should be discovered by recursively walking a build folder.
We keep an explicit list for important static pages such as the homepage, about page, contact page, privacy policy, and other one-off routes.
That lets us assign the right source files, update frequency, and priority instead of pretending every URL has the same role.
The sitemap becomes a deliberate representation of the site rather than a dump of every file Next.js happened to generate.
Dynamic page families come from content
The more interesting pages are generated from content systems.
A service detail route may exist once in code but render six service pages. A location route may render a dozen locations. The article route may render dozens of MDX files.
The sitemap should represent the rendered pages, not the number of route templates.
That means the content loader becomes part of search infrastructure.
This is one reason we care about clean content models. A good content system does more than render text. It can power navigation, metadata, related links, static generation, and sitemap entries from the same source of truth.
We escape XML instead of assuming content is safe
Generated XML still has syntax rules. URLs and values can contain characters that need escaping.
Our sitemap routes explicitly escape ampersands, quotes, apostrophes, angle brackets, and other XML-sensitive characters before building the response.
That is a small implementation detail, but it is the kind of thing that separates "we have a sitemap" from "we built a sitemap generator."
Why we sometimes use a route instead of Next.js metadata helpers
Next.js can generate sitemaps through its metadata conventions, and we use that approach on simpler sites.
On another project, the sitemap is a straightforward array of known routes returned through the framework's sitemap type. That is a good fit because the site has a small, controlled page set.
On more content-heavy sites, we use a custom XML route because we want more control over source-derived lastmod, XSL presentation, mixed content families, and exact output.
The choice follows the complexity of the site.
The sitemap is not an SEO magic trick
A sitemap does not make a weak page rank. It does not substitute for internal links. It does not turn duplicate pages into distinct search assets.
Its job is more modest and more important: provide search engines with a reliable machine-readable inventory of canonical public URLs and useful recrawl hints.
We want that inventory to be accurate every time the site builds.
That is why our sitemap is usually downstream of the content architecture instead of being a separate artifact someone has to remember to update.
From explanation to proof
Where this connects to the work
This category connects our search recommendations to implementation details that can be inspected, tested, and maintained.