Staging Sites That Search Should Not Index

By Simeon Matheka, Founder & Creative Director · Published 2026-09-28 · Updated 2026-09-28 · 13 min read

A public staging host is a second site. Keep it off Google with a password or with noindex that the crawler can actually see, on a hostname you do not submit. Strip that block before production. DNS cutover stays in the migration runbook.

A side door with a paper band reading staging and a noindex slip, beside an open front door for the live host

The staging link is in a slide, a ticket, and a footer someone forgot to change. It is a real hostname with real HTML. If Google can find it, you now have two copies of the site. One of them is the draft.

DNS, the redirect map, and rollback belong to the migration runbook. Stay here for one job: the staging host does not get indexed. When you move the live domain, follow that runbook and come back only to confirm the block is still on staging and gone from production.

Staging is a different host

Framework. Put the draft on its own hostname. A subdomain or a separate host is enough, as long as it is not the hostname you want in Search. The leak is then a URL you can point at, password, or remove. A draft that shares the production hostname, with a secret path you hope nobody links, is how a client PDF becomes an indexable URL.

The production files, on the host you do mean to index, are a different setup. For a static site on this stack, that lives in the Cloudflare hosting guide. Do not reuse the production hostname as a convenience for reviewers. Reviewers can use the staging host. Google should be aimed at the other one.

A public staging host is also an access problem. Draft copy, unused forms, and half-finished admin paths do not belong on the open web. The threat model for a marketing site is the security guide. This page only covers the index. Lock the host for both reasons, and do not pretend a robots file did the locking.

What robots.txt is for

Sourced from Introduction to robots.txt (Search Central, last updated 2025-12-10 UTC, checked 28 September 2026). A robots.txt file tells search engine crawlers which URLs they can access on your site. Search Central says this is used mainly to avoid overloading your site with requests. It is not a mechanism for keeping a web page out of Google.

The same page is blunt about the failure. Do not use robots.txt to hide web pages from Google Search results. If other pages point at your URL with descriptive text, Google could still index the URL without visiting the page. A blocked URL can still appear in results, and the result will not have a description. Google says it will not crawl or index the content blocked by robots.txt, and it might still find and index a disallowed URL that is linked from elsewhere. The URL, and anchor text from those links, can show up.

robots.txt is also a request, not a lock. Search Central says the instructions cannot force every crawler to obey. Googlebot and other respectful crawlers do. Others might not. Rules differ by crawler. If you need the draft private, password-protect the files on the server. That is the sentence Search Central uses for keeping information secure, and it is one of the methods it names for keeping a URL out of Google Search results. The others on that page are noindex, or removing the page.

If you use a content management system (CMS), you might not edit robots.txt by hand. Search Central says the CMS might expose a search setting instead. Look up how that product changes page visibility. Do not assume the setting writes the rule you think it writes. Check the response.

noindex works only when the crawler can read it

Sourced from Block Search indexing with noindex (Search Central, last updated 2025-12-10 UTC, checked 28 September 2026). noindex is a meta tag or an HTTP response header. Search engines that support it, including Google, use it to keep that content out of the index. When Googlebot crawls the page and extracts the tag or header, Google drops that page entirely from Google Search results, even if other sites link to it.

There is a condition on that sentence. For noindex to work, the page must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If robots.txt blocks the URL, or the crawler cannot access the page, the crawler never sees noindex. The page can still appear in search results, for example if other pages link to it. That is the staging trap. Disallow in robots.txt stops the crawl. A page Google cannot crawl cannot show Google the noindex rule. You get the URL in results, often with no description, and the tag you were counting on never ran.

Search Central also says specifying noindex inside the robots.txt file is not supported by Google. Put the rule on the response. Two forms, same effect. Pick the one that matches the file type.

  • HTML meta tag: In the head, a robots meta tag with content noindex. A googlebot meta tag with content noindex limits the rule to Google’s web crawlers. Other engines may interpret noindex differently, so a page can still appear outside Google.
  • HTTP header: X-Robots-Tag: noindex, or none. Search Central says a response header fits non-HTML files such as PDFs, video, and images. Use it when the staging host serves those files too.
<meta name="robots" content="noindex">

You can combine noindex with other indexing rules. Search Central’s example is noindex together with nofollow on the same meta tag. Combining crawl rules and index rules can also cancel the result you wanted. The robots.txt introduction warns that stacking rules can make them counteract each other. The noindex page is the specific case: a robots block hides the noindex you added.

Password protection sits in both documents, and they do not collapse into one slogan. The robots.txt introduction lists password protection as a way to keep a URL out of Google Search results, and as the better option when you need crawlers kept off private files. The noindex page says that if the crawler cannot access the page, it will never see noindex, and the URL can still appear when something else links to it. Read that as two tools with different jobs. A password is how you keep the draft private, which is what you want for staging. noindex is how you tell a crawler that was able to fetch the page to drop it. Do not put noindex behind a wall and call the wall a substitute for the tag. Do not put a Disallow in front of noindex either.

ControlWhat Search Central saysStaging use
robots.txt disallowManages crawl access. Not a way to keep a page out of Google. The URL can still appear, without a description.Do not use this alone, and do not put it in front of noindex.
noindex meta tag or X-Robots-TagOnce crawled and seen, Google drops the page from Search results even if other sites link to it. The crawler must be able to fetch the page.Use on staging only if the crawler can see it. Leave it off production.
noindex written inside robots.txtNot supported by Google.Do not.
Password on the serverNamed as a way to keep a URL out of Google, and as the way to keep files private from crawlers that ignore robots.txt.Use on the staging host so the draft is not public.

Do not hand Google the staging host

Framework. The sitemap and the Search Console property are how you tell Google which host you claim. Do not submit a sitemap of staging URLs. Do not list the staging hostname inside the sitemap you submit for the live site. Do not verify the staging host and then treat it as the property you want in the index. The rule is which host you offer.

Links leak anyway. A public doc, a partner site, or an old campaign URL can point at staging. That is the case both Google pages describe: a URL discovered from other pages. Password protection is the control that does not depend on Google obeying a text file. noindex is the control that drops a URL Google was able to crawl. Use the password on staging. Add noindex on responses the crawler is actually allowed to fetch. Keep robots.txt from blocking that fetch if noindex is the plan.

If the staging app is a JavaScript shell, the first HTML may be empty even when a person sees a page. That failure mode is why crawlers saw an empty React site. An indexed staging URL that is also an empty shell is two problems. Fix the shell on the host you intend to index. Keep the staging host out of the index either way.

Strip the block before production

The dangerous moment is the copy from staging to the live response. noindex is doing its job on staging. The same bytes on production tell Google to drop the real pages once they are crawled. Search Central says that drop happens regardless of other sites linking to you. A forgotten meta tag is enough.

  1. Production response: Fetch the live URL you are about to announce. The HTML head should not contain a noindex robots meta tag. The response should not send X-Robots-Tag: noindex or none.
  2. Production robots.txt: It should not disallow the pages you want crawled. A disallow left over from staging will not “protect” those URLs. It can leave them eligible to appear without a description, and it will hide a noindex you might still be relying on.
  3. Staging response: The password stays. If you also rely on noindex, the crawler has to be able to see it. Confirm you did not swap the rules during the copy.
  4. Sitemaps: The sitemap submitted for the live host lists live URLs only. Staging URLs are absent.
  5. If a staging URL is already showing: Search Central says Google has to crawl the page to see the tag. If you just added noindex, the URL can linger because Googlebot has not been back. That wait can be months. The URL Inspection tool can request a recrawl. If robots.txt is the reason the tag was invisible, edit robots.txt so Google can fetch the page. The Page Indexing report shows pages where Googlebot extracted noindex. For a fast removal, Search Central points at its removals documentation. Use that path when a draft URL is already public in results.

Then stop. The hostname change, the redirects, and the rollback plan are the migration runbook. Follow that list for the move.

Hypothetical: Disallow plus a public link

Hypothetical, labeled as such. A team puts the new site on staging.example.com. They add a robots.txt that disallows the whole host. They also add a noindex meta tag, because someone said to do both. A partner page on the public web links to a staging case study with a descriptive anchor. Google cannot crawl the URL, so it never reads noindex. The URL can still appear, without a description. The team discovers it from a branded query, not from a lab.

The repair that matches the docs: remove the disallow that hides the tag if you need Google to see noindex, or password-protect the host so the draft is not a public URL in the first place. Prefer the password for a staging server. Confirm the live host never received the meta tag. Do not “fix” it by submitting a staging sitemap. That offers the wrong host.

Put the check on the maintenance list

Once a month, fetch one staging URL and one production URL. Staging asks for a password, or it shows noindex to a crawler that is allowed to look. Production shows neither the meta tag nor the header. The maintenance blueprint is the home for that recurring check. Access control beyond the index, including who can reach draft tools, stays with the security guide.

What this does not prove: that every other search engine will drop the URL the way Google describes. Search Central says some engines interpret noindex differently, and that not every crawler obeys robots.txt. Password protection is the control that does not depend on that cooperation. A clean production response is the control that stops you from noindexing yourself.

If the draft and the live host are still the same URL, split them as part of the website work before launch week. Start a conversation with the staging hostname and a copy of the production response headers, not with a DNS plan. The DNS plan has its own page.

Frequently asked questions

Is a robots.txt disallow enough to keep staging out of Google?

No. Search Central’s introduction to robots.txt, checked 28 September 2026, says a robots.txt file tells crawlers which URLs they can access, mainly to avoid overloading the site. It is not a mechanism for keeping a web page out of Google. A disallowed URL can still be indexed if other sites link to it. The result may show the URL without a description. To keep a page out of Google, Search Central points you to noindex or password protection.

Can I put noindex inside robots.txt?

Google does not support that. The noindex documentation says you implement noindex as a meta tag or as an HTTP response header. Specifying the noindex rule in the robots.txt file is not supported. A content management system (CMS) may offer a setting instead of a raw file. Use the setting or the meta tag or the header. Do not invent a robots.txt line and call it noindex.

Why would noindex fail if robots.txt also blocks the URL?

Because Googlebot has to crawl the page to see the rule. Search Central says the page must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If robots.txt blocks the URL, or the crawler cannot access the page, the crawler never sees the noindex rule. The page can still appear in search results, for example if other pages link to it. Disallow plus noindex is how staging URLs leak with no description and no drop.

Should I submit a sitemap for the staging host?

No. Framework: do not submit staging URLs, and do not add the staging hostname to the sitemap you submit for the live site. This page does not claim a crawl-priority number for sitemaps. The operating rule is simpler. The only host you hand to Google as your site is the one you want indexed. Cutover of the live domain is a different job.

When do I remove noindex?

Before the response people should find is the one Google fetches. Once Googlebot crawls a page and extracts noindex, Search Central says Google drops that page from Google Search results, even if other sites link to it. Leave noindex on the staging host. Remove the meta tag and any X-Robots-Tag: noindex from the production response. Confirm production robots.txt is not blocking the URLs you want crawled. DNS, redirects, and rollback stay in the migration runbook.

Tags: staging, noindex, robots.txt, indexing, Search Console, website maintenance