Back to BlogTroubleshooting

Next.js Sitemap and Canonicals Pointing at the Wrong Domain: The Fix

Rupak Amin

Founder & Lead Engineer, RAITHub

10 min read

If your Next.js sitemap, robots.txt and canonical tags point at the wrong domain, the cause is almost always the value they are built from, usually NEXT_PUBLIC_SITE_URL left on an old host. Correct it and redeploy, because NEXT_PUBLIC_ values are fixed at build time. Then make the code reject any host that is not yours, so a stale value cannot reach indexed markup again.

This happened on the RAITHub website. The production value of NEXT_PUBLIC_SITE_URL still pointed at an old Vercel alias that now returns 404. Every canonical tag on www.raithub.com, the sitemap and the robots.txt sitemap line all pointed at that dead host. The pages themselves worked perfectly, which is why it went unnoticed until an SEO audit read the live sitemap. This post covers why it happens, how to check your site, and the fix we shipped.

Why does a wrong domain in the sitemap and canonicals matter?

Because you are telling search engines that the real copy of every page lives somewhere else. Google describes rel="canonical" as "a strong signal that the specified URL should become canonical", and a sitemap as a weaker one (Google Search Central: consolidate duplicate URLs).

When the canonical points at a host that 404s, your correct page is declaring a broken page as its original. Google's guidance also says not to give different canonical URLs for the same page through different methods, such as one URL in the sitemap and another in rel="canonical". A half-fixed site, with correct canonicals but a stale sitemap, sends exactly that mixed signal.

The pages still load for visitors, nothing errors, and no test fails. You only see it by reading the markup or by watching indexing reports stall.

Why does NEXT_PUBLIC_SITE_URL go stale?

Usually because the project changed domain, or was renamed, and one environment variable in one environment was never updated. Several details of Next.js and Vercel make that easy to miss.

CauseWhat happensSource
NEXT_PUBLIC_ values are inlined at buildChanging the variable does nothing until you rebuild; old deployments keep the old valueNext.js: environment variables
Values are set per environmentProduction can keep an old host while Preview and Development look correctVercel project settings
VERCEL_URL used as the site URLIt is the generated deployment host, such as *.vercel.app, with no https://Vercel: system environment variables
VERCEL_PROJECT_PRODUCTION_URL used before a custom domain existsVercel picks the shortest production custom domain, or a vercel.app domain if there is noneVercel: system environment variables
Each file reads the variable itselfFixing it in the layout leaves the sitemap, robots or JSON-LD wrongThis site, before the fix

The Next.js docs are explicit on the first point: after being built, "your app will no longer respond to changes to these environment variables", and NEXT_PUBLIC_ values "will be frozen with the value evaluated at build time". Fixing the variable in the dashboard without redeploying changes nothing.

On this site, the last row was the real weakness. The layout, sitemap, robots route and structured data each read process.env.NEXT_PUBLIC_SITE_URL directly, so one bad value spread everywhere at once.

How do you check which domain your sitemap and canonicals use?

Read the production output directly. Three commands cover the sitemap, robots.txt and a page's canonical tag:

# Hosts listed in the sitemap (should be one line: your domain)
curl -s https://www.example.com/sitemap.xml | grep -o 'https://[^/]*' | sort | uniq -c

# The Sitemap and Host lines in robots.txt
curl -s https://www.example.com/robots.txt | grep -iE 'sitemap|host'

# The canonical tag and og:url on a real page
curl -s https://www.example.com/blog/some-post | grep -oE '<link rel="canonical"[^>]*>|property="og:url"[^>]*'

Every URL should use your production domain, with the same protocol and the same www or apex choice as your redirects. If any line shows a vercel.app host, a preview URL or localhost, the value used at build time was wrong.

Also check a preview deployment. Its robots.txt should block crawling, and its canonicals should still point at production, not at itself.

What was the fix on this site?

We made one module the only source of the public site URL, and made it refuse any host that is not ours. The environment variable is still read, but it can only choose between raithub.com hosts.

// src/lib/site-url.ts
export const CANONICAL_SITE_URL = 'https://www.raithub.com'

const ALLOWED_HOST = /(^|\.)raithub\.com$/i

export function resolveSiteUrl(value: string | undefined): string {
  if (!value) return CANONICAL_SITE_URL
  try {
    const url = new URL(value.trim())
    if (url.protocol !== 'https:' || !ALLOWED_HOST.test(url.hostname)) return CANONICAL_SITE_URL
    return url.origin
  } catch {
    return CANONICAL_SITE_URL
  }
}

/** Absolute origin without a trailing slash, e.g. https://www.raithub.com */
export const SITE_URL = resolveSiteUrl(process.env.NEXT_PUBLIC_SITE_URL)

Four details matter:

  • An allow-list, not a block-list. We do not try to spot "bad" hosts. Anything that is not raithub.com, or a subdomain of it, falls back to the canonical domain.
  • The pattern is anchored. It accepts www.raithub.com and staging.raithub.com, but rejects look-alikes such as raithub.com.evil.example and notraithub.com. The unit tests assert both.
  • https only. An http:// value is treated as a mistake and falls back too.
  • Origin only. A value with a trailing slash or a path is reduced to its origin, so joined URLs never get a double slash.

Every consumer now imports SITE_URL instead of reading the environment. The root layout sets metadataBase: new URL(SITE_URL), so relative canonicals resolve against the allowed host. Next.js documents that metadataBase lets URL-based metadata fields use a relative path, composed into a fully qualified URL (Next.js: generateMetadata, metadataBase). The sitemap and robots routes use the same constant:

// src/app/robots.ts (trimmed; the real file also lists AI crawlers)
import { SITE_URL } from '@/lib/site-url'

export default function robots(): MetadataRoute.Robots {
  // Preview and development deploys are never crawled.
  if (process.env.VERCEL_ENV === 'preview' || process.env.VERCEL_ENV === 'development') {
    return { rules: [{ userAgent: '*', disallow: '/' }] }
  }
  return {
    rules: [{ userAgent: '*', allow: '/', disallow: ['/admin/', '/api/'] }],
    sitemap: SITE_URL + '/sitemap.xml',
    host: SITE_URL,
  }
}

Why warn about the stale value instead of failing the build?

Because once the allow-list exists, a stale value no longer breaks indexed markup, and failing would block a deploy that is otherwise correct. The build still says something is off.

The production environment check runs from next.config.ts on production deploys only. It already fails the build when a required variable is missing or is not an https URL. For the host, it warns:

// src/lib/env.ts (trimmed; the real message is longer)
const staleHosts = ['NEXTAUTH_URL', 'NEXT_PUBLIC_SITE_URL'].filter((n) => {
  const value = process.env[n]
  if (!value) return false
  try {
    return !/(^|\.)raithub\.com$/i.test(new URL(value).hostname)
  } catch {
    return false
  }
})
if (staleHosts.length > 0) {
  console.warn(staleHosts.join(', ') + ' not on raithub.com. Set to https://www.raithub.com in Vercel.')
}

The warning also covers NEXTAUTH_URL, which the allow-list does not protect. Auth callbacks use it as set, so a stale value there still needs fixing in the project settings. The warning is the prompt to do it.

What should you do after fixing the domain?

Redeploy, verify, then tell search engines. In order:

  1. Set the variable correctly in the Production environment of your host, and redeploy so the new value is built in.
  2. Re-run the three curl checks on the live site. All hosts should now match.
  3. Resubmit the sitemap in Google Search Console and Bing Webmaster Tools, and use URL inspection on a few important pages to request recrawling.
  4. Redirect the old host if you still control it. A permanent redirect from every old path to the same path on your domain passes on links and bookmarks. If the old host is gone, the corrected canonicals do the work, more slowly.
  5. Keep preview deployments out of the index. Vercel does not index preview deployments by default, using a noindex header (Vercel: are preview deployments indexed?). A robots rule like the one above is a second layer.

Expect the corrections to take effect over weeks as pages are recrawled, not overnight. We make no promise about how long a given search engine takes.

Why RAITHub for this?

Because this is a bug we shipped, found and fixed on our own production site, with the fix written so it cannot happen the same way again. Technical SEO failures like this are configuration bugs, and they need the same treatment as code bugs: find the single source of truth, test it, and add a guard.

  • One source of truth. Site URLs, canonicals and structured data come from one tested module.
  • Tests on the rule. The allow-list has unit tests for unset values, look-alike hosts, http and trailing slashes.
  • Checks on what ships. We verify with curl on the live site, not by reading the code. See the Next.js production checklist for the rest of the launch checks, or our Next.js development service.

When you don't need us

  • The variable is simply wrong in one environment. Fix it, redeploy, run the curl checks and resubmit the sitemap. That is an hour's work.
  • You only have one domain and a hard-coded constant. A constant cannot go stale. Keep it, and make every file import it.
  • The problem is ranking, not markup. If your canonicals are right and pages still do not rank, that is a content and authority question, not an engineering one.

If canonicals are still wrong after a redeploy, see how a fixed-price fix works, or send us your sitemap URL for a free 15-minute technical audit.

Tested on Next.js 16 (App Router) deployed on Vercel. Last reviewed: 29 September 2026.

Frequently asked questions

Why does my Next.js sitemap show the wrong domain?

The sitemap builds absolute URLs from a base URL, usually NEXT_PUBLIC_SITE_URL or a Vercel system variable. If that value was wrong when the app was built, every URL in the sitemap uses the wrong host. Fix the value and redeploy.

Do I need to redeploy after changing NEXT_PUBLIC_SITE_URL?

Yes. Next.js inlines NEXT_PUBLIC_ values at build time, and the docs state that a built app no longer responds to changes in them. The new value only takes effect in a new build.

Should I use VERCEL_URL for canonical URLs?

No. VERCEL_URL is the generated host of each deployment, so every deploy would declare a different canonical. Use your production domain from one constant or a validated variable.

What is metadataBase in Next.js?

A base URL set in the root layout that turns relative metadata URLs, such as alternates.canonical: '/blog/x', into absolute ones. If it points at the wrong host, every relative canonical inherits the mistake.

How long does Google take to pick up corrected canonicals?

It depends on how often your pages are crawled. Resubmitting the sitemap and requesting indexing for key pages helps, but expect it to happen gradually over weeks rather than immediately.

Should preview deployments have their own canonical URLs?

No. Previews should point canonicals at production and block crawling in robots.txt. Vercel also sends a noindex header on preview deployments by default.

Next.jsSitemapCanonical URLrobots.txtNEXT_PUBLIC_SITE_URLVercelTechnical SEO

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.