Skip to Content

Scrunch AI Platform, Explained: What Your Site Actually Needs

Jill Roberson

 Sitecore's acquisition of Scrunch changes the AI-visibility conversation, but not the infrastructure it takes to win it. 

Customers increasingly ask ChatGPT, Gemini, Perplexity, and Claude which vendor or organization is worth talking to before they ever open a browser tab. Those same AI tools then send out agents like GPTBot, ClaudeBot, PerplexityBot, and others to actually crawl and read websites while researching the answer on a person's behalf. Scrunch was built around that shift: part of it watches what AI says about your brand, and part of it watches (and shapes) what AI agents actually find when they show up at your site.

Which means the agents are already on your site. The real question is whether your infrastructure can do anything about it.

Scrunch is one of the fastest-growing platforms in AI search visibility, and its own site claims it's already trusted by more than 500 companies. But "Scrunch" isn't one product, it's a platform made up of several components, and they don't all work the same way. We're going to focus on the big three:

  • Monitoring & Citations
  • Agent Traffic
  • Agent Experience Platform (AXP)

Monitoring & Citations runs the moment you turn it on. Agent Traffic and AXP need a little more tinkering - both require a front-end host you actually control before they can do anything. Here's why, and how Dataweavers' Arc and Spark are built to close that gap.

The three core components of Scrunch 

1. Monitoring & Citations: works the moment you switch it on

Monitoring & Citations runs prompts against the major AI platforms including ChatGPT, Gemini, Perplexity, Google AI Overviews, using a mix of browser automation and official APIs, then reports back on whether and how your brand shows up, which sources are shaping those answers, and how you stack up against competitors.

None of that touches your website. Scrunch is asking the AI platforms questions and recording the answers, the same way a human analyst would, just at scale. There's no code to install, no server to configure, no CMS integration required. Point it at your brand and your target prompts, and it starts reporting the same day. That's what "out of the box" actually means here: the work happens entirely on Scrunch's side of the fence.

2. Agent Traffic: needs to see what's hitting your server

Agent Traffic is a different kind of data entirely. It separates AI bot visits by type - training, indexing, and retrieval - and shows which of your pages agents actually rely on when generating an answer about your brand.

That information doesn't come from asking an AI model anything. It comes from watching real requests arrive at your website. And requests only exist at one layer: your edge network or CDN i.e. the layer of servers that decides what gets served before a request ever reaches your origin. To read that traffic, Scrunch has to be instrumented into that layer. If your delivery layer is a black box, bundled inside a vendor platform, cobbled together under deadline, or simply not yours to configure, there's nothing for Scrunch to plug into, and Agent Traffic has nothing to report.

3. AXP: needs to serve a different response, on the fly

AXP goes a step further than watching. When a request comes in, it has to:

  1. Check whether the request is coming from a human or an AI agent
  2. If it's an agent, serve a token-light, structured version of that page built for machine reading
  3. If it's a human, serve the normal site, completely untouched

Same URL, two different responses, decided in real time. Scrunch's own numbers put the payoff at roughly a 96% token reduction on an optimized page, from over 115,000 tokens down to under 5,000. But that's not a content trick; it's a hosting-layer capability. Rewriting what gets returned for a given request, based on who's asking, can only happen at the point where your site actually fulfills that request — which means AXP needs exactly what Agent Traffic needs, just one level deeper: a front-end host that can be configured to intercept and rewrite responses at the edge.

Why "out of the box" stops after one capability 

The dividing line across Scrunch's platform isn't features, it's where the work happens.

  • Monitoring & Citations happens on Scrunch's side, watching AI platforms from the outside.
  • Agent Traffic and AXP happen on your side, reading and rewriting what your own website hands back. The moment Scrunch needs to touch your delivery layer, that layer has to be touchable: observable, configurable, and genuinely yours to instrument.

A lot of enterprise sites that moved to headless don't have that. Rendering hosts get stitched together under a deadline, edge logic becomes something nobody wants to touch, and hosting ends up locked inside someone else's platform. Monitoring & Citations doesn't care. Agent Traffic and AXP absolutely do.

This is exactly what Arc by Dataweavers is built for 

Arc by Dataweavers is a self-service, Azure-native front-end hosting platform for headless Sitecore and Sitecore AI experiences, deployed inside your own Azure tenant rather than bolted onto someone else's. That means the delivery layer - routing, edge logic, environments, governance - is yours from day one, not something your team has to fight to expose later.

That's precisely the condition Agent Traffic and AXP require. A site running on Arc already has a delivery layer that's open, governed, and configurable, the layer Agent Traffic reads from, and the layer AXP would serve through. Instead of a separate engineering project just to expose bot detection or edge routing, teams building on Arc already have the foundation those two products need.

That's what it means to be Scrunch-ready: not a certification, but an architecture where the hardest technical prerequisite for AI agent traffic and content delivery is already solved before Scrunch ever enters the picture.

Spark: making sure the agent sees the current version of you 

None of this matters if the content underneath is stale. That's where Spark comes in.

Many headless Next.js setups rely on timeout-based incremental static regeneration (ISR), the page only refreshes after a set amount of time, or after enough visitor traffic triggers a rebuild. That delay is invisible to most teams, right up until it isn't. Spark replaces it with on-demand ISR: the moment content is published in Sitecore Experience Edge, Spark triggers an immediate, targeted regeneration of that page, no waiting on a timer, no waiting on traffic.

For Scrunch, that alignment is the whole game. When someone reads an AI answer that cites your brand and clicks through, they expect to land on a page that matches what they just read. Scrunch's own Sitecore integration keeps the AI-facing side of AXP current, so that part isn't in question. What's still on you is the human-facing page: if your site is serving actual visitors a version that's behind what Scrunch already has synced for AI, the two experiences drift apart, and that gap, not the AI summary itself, is what erodes trust when someone clicks through expecting one thing and finds another.

Spark closes that gap by keeping publishing near-instant, so the human page a visitor lands on stays in step with whatever's already fresh on the AI side. It's available as a standalone product for any Sitecore Experience Edge or Next.js setup, and it ships as a built-in feature of Arc, so Arc customers get instant publishing without adding a separate integration.

The short version

  • Monitoring & Citations works out of the box because it watches AI platforms from the outside, no hosting dependency at all.
  • Agent Traffic and AXP work at your site's delivery layer, so they need a front-end host that's open enough to instrument and configure.
  • Arc gives you that host by design, an Azure-native delivery layer you control, ready for the parts of Scrunch that need one.
  • Spark keeps the content behind all of it current, so what AI agents retrieve (and cite) is never stale and when humans hit cited links they get to the right version of a page.

Where to go from here 

If you're evaluating Scrunch and you're running on Sitecore, Monitoring & Citations will work on day one regardless of your hosting, that part's easy. The real question is whether your front-end host can support Agent Traffic and AXP when you're ready to turn them on. Arc is built so the answer is already yes, and Spark makes sure whatever Scrunch finds or serves is actually current.  

Frequently Asked Questions

What are the three parts of the Scrunch platform?

Scrunch is built around several features but three core capabilities: Monitoring & Citations, which tracks how your brand appears in AI answers; Agent Traffic, which analyzes AI bot traffic hitting your site; and AXP (Agent Experience Platform), which detects AI agents at the edge and serves them an optimized version of your pages.

Do you need special hosting to use Scrunch?

Not for Monitoring & Citations, it runs independently of your website by querying AI platforms like ChatGPT and Perplexity directly. Agent Traffic and AXP are different: both are directed by your sites delivery layer/CDN, so they need a front-end host that lets Scrunch read what's hitting your site (Agent Traffic) or rewrite what gets served to a given request (AXP).

What does AXP actually do to my website?

AXP checks each incoming request to see whether it's coming from an AI agent. If so, it serves a token-light, structured version of that page built for machine reading; human visitors to the same URL get the normal site, untouched. By Scrunch's own figures, that can shrink a page from well over 100,000 tokens to under 5,000.

Does Spark only work with Arc by Dataweavers?

No. Spark is a standalone product that delivers on-demand ISR for any Sitecore Experience Edge or Next.js setup. It's also built directly into Arc as a feature, so Arc customers get instant publishing without adding a separate integration.

Why does content freshness matter for AI search visibility?

Because Scrunch's Sitecore integration already keeps the AI-facing side current, the real risk sits on your side: the human page. If someone reads an AI answer that cites your brand and clicks through expecting to see what they just read - current pricing, live messaging, a product that still exists - but lands on a page that hasn't caught up yet, that mismatch is what damages trust, not the AI citation itself. Spark closes that gap by keeping publishing near-instant, so the page a human visitor lands on stays in sync with whatever Scrunch already has synced on the AI side.