opdeck / blog / how-fast-is-the-web-explore-billions-of-real-user-measuremen

How to Measure Web Speed Using Billions of Real-User Data with BEACON

September 29, 2026 / OpDeck Team
Web PerformanceReal User MonitoringSpeed MeasurementBEACONUser Data

What Real User Monitoring Actually Tells You (And How to Act on It)

The web performance conversation has long been dominated by synthetic benchmarks — automated tools running Lighthouse audits in controlled environments, reporting scores that may bear little resemblance to what your actual users experience. Cloudflare's release of the BEACON dataset changes that conversation significantly. By open-sourcing billions of anonymized Real User Monitoring (RUM) records on Google BigQuery, Cloudflare has given developers a rare window into how the web actually performs across browsers, regions, and device types in the wild.

But understanding what that data means — and more importantly, knowing what to do about it — requires a solid grasp of the metrics involved, the methodology behind RUM, and how to apply these insights to your own sites. This guide breaks all of that down.


Understanding Real User Monitoring vs. Synthetic Testing

Before diving into the BEACON dataset itself, it's worth clarifying the fundamental difference between RUM and synthetic monitoring, because they serve different purposes and answer different questions.

Synthetic testing runs predefined scripts against your site from controlled environments. Tools like Lighthouse, WebPageTest, and OpDeck's Website Performance Analyzer fall into this category. They're reproducible, comparable across time, and excellent for catching regressions during development. You control the conditions: network throttling, device emulation, geographic origin.

Real User Monitoring instruments your actual production traffic. JavaScript snippets collect timing data from real browsers on real networks and devices, then send that telemetry back to a collection endpoint. The result is a distribution of experiences rather than a single score — and that distribution tells a very different story.

Consider a typical scenario: your Lighthouse score is 92, but your 75th percentile Largest Contentful Paint (LCP) in production is 4.2 seconds. Synthetic tests run on a fast connection from a US data center. Your users include people on 3G networks in Southeast Asia, older Android devices in rural areas, and corporate users behind content-filtering proxies. RUM captures all of them.

The BEACON dataset aggregates this kind of real-world telemetry at a scale most organizations could never achieve individually — billions of records across countless sites, browsers, and regions. Analyzing it reveals what "normal" looks like on the web, which is invaluable context when benchmarking your own performance.


Core Web Vitals: The Metrics That Actually Matter

The BEACON dataset focuses heavily on Core Web Vitals, Google's standardized set of user-centric performance metrics. Understanding what each metric measures — and what causes it to degrade — is essential before you can act on any dataset.

Largest Contentful Paint (LCP)

LCP measures when the largest visible content element (typically a hero image, video poster, or large text block) finishes rendering. Google's threshold for "Good" is under 2.5 seconds; "Needs Improvement" is 2.5–4.0 seconds; anything above 4.0 seconds is "Poor."

Common LCP culprits:

  • Render-blocking resources (CSS and JavaScript that delay the browser's rendering pipeline)
  • Slow server response times (Time to First Byte above 600ms)
  • Images without explicit dimensions causing layout recalculations
  • Missing fetchpriority="high" on the LCP image element
  • No preload hints for critical resources
<!-- Correct: preload your LCP image -->
<link rel="preload" as="image" href="/hero.webp" fetchpriority="high">

<!-- Also add fetchpriority directly on the img element -->
<img src="/hero.webp" fetchpriority="high" alt="Hero image" width="1200" height="600">

Interaction to Next Paint (INP)

INP replaced First Input Delay (FID) as the interactivity Core Web Vital in 2024. Where FID only measured the delay before the browser could respond to the first interaction, INP measures the worst interaction latency throughout the entire page lifecycle. Good is under 200ms; Poor is above 500ms.

INP failures almost always come down to long tasks on the main thread:

// Bad: synchronous processing that blocks the main thread
function processLargeDataset(data) {
  return data.map(item => heavyTransformation(item));
}

// Better: break work into chunks using scheduler.postTask or setTimeout
async function processLargeDatasetAsync(data) {
  const results = [];
  const chunkSize = 50;
  
  for (let i = 0; i < data.length; i += chunkSize) {
    const chunk = data.slice(i, i + chunkSize);
    results.push(...chunk.map(item => heavyTransformation(item)));
    
    // Yield to the browser between chunks
    await new Promise(resolve => setTimeout(resolve, 0));
  }
  
  return results;
}

Cumulative Layout Shift (CLS)

CLS quantifies unexpected layout movement — elements jumping around as the page loads. Good is below 0.1; Poor is above 0.25.

The most common CLS sources:

  • Images and videos without explicit width and height attributes
  • Dynamically injected content (ads, cookie banners, chat widgets)
  • Web fonts causing text reflow (FOIT/FOUT)
  • Animations that modify layout properties instead of using transform
/* Bad: animating layout properties causes CLS */
.expanding-element {
  transition: height 0.3s ease;
}

/* Good: use transform which doesn't trigger layout */
.expanding-element {
  transition: transform 0.3s ease;
  transform-origin: top;
}

What the BEACON Data Reveals About Global Performance

The BEACON dataset's value lies in its geographic and browser diversity. When you query it on BigQuery, patterns emerge that synthetic testing would never surface.

Regional Disparities Are Larger Than Most Developers Assume

Developers in North America and Western Europe building for global audiences routinely underestimate how different the experience is for users elsewhere. The BEACON data confirms what performance engineers have long suspected: median LCP in regions with lower average connection speeds can be 3–5x slower than in high-connectivity markets.

This has direct implications for infrastructure decisions:

  • CDN edge node coverage matters enormously. A CDN with strong presence in Southeast Asia, South America, and Africa will dramatically outperform one optimized for US/EU traffic.
  • Image optimization is not optional for global audiences. Serving WebP or AVIF with appropriate compression can cut image payload by 50–80%.
  • Third-party scripts are disproportionately painful on slower connections. A 200KB analytics bundle that loads in 100ms on fiber takes over a second on 3G.

Browser Differences Are Significant

Chromium-based browsers dominate the dataset, but Safari and Firefox show meaningfully different performance profiles for the same sites. This is partly due to different JavaScript engine optimizations, partly due to device correlations (Safari users skew toward iOS/macOS hardware), and partly due to feature support differences affecting how polyfills load.

Safari's lack of support for certain modern APIs means sites serving polyfills conditionally need robust feature detection:

// Feature detection before loading polyfills
if (!('IntersectionObserver' in window)) {
  import('/polyfills/intersection-observer.js').then(() => {
    initializeLazyLoading();
  });
} else {
  initializeLazyLoading();
}

Soft Navigation Metrics: The SPA Challenge

One of the more interesting aspects of the BEACON dataset is its inclusion of soft navigation metrics — a relatively new area of measurement designed to capture performance in Single Page Applications (SPAs) and frameworks like Next.js, Nuxt, and SvelteKit.

Traditional Core Web Vitals are measured from navigation start (hard page loads). But in an SPA, users rarely trigger full page loads after the initial one. Client-side routing changes the URL and updates the DOM, but the browser doesn't register a new navigation. This means LCP, CLS, and INP measurements after the first page load are invisible to traditional RUM instrumentation.

The Soft Navigations API (currently experimental in Chrome) attempts to solve this by detecting "navigation-like" DOM and URL changes:

// Observing soft navigations with PerformanceObserver
const observer = new PerformanceObserver((list) => {
  for (const entry of list.getEntries()) {
    if (entry.entryType === 'soft-navigation') {
      console.log('Soft navigation detected:', entry.name);
      console.log('Duration:', entry.duration);
    }
  }
});

observer.observe({ type: 'soft-navigation', buffered: true });

The BEACON data's inclusion of soft navigation metrics at scale helps establish what "good" looks like for SPA performance — previously an area with almost no real-world benchmarks.


Querying the BEACON Dataset: Getting Started

The dataset lives on Google BigQuery under Cloudflare's public project. To start exploring it, you'll need a Google Cloud account (BigQuery has a generous free tier for queries under 1TB/month).

Basic Query Structure

-- Find median LCP by country for the past 30 days
SELECT
  country_code,
  APPROX_QUANTILES(lcp, 100)[OFFSET(50)] AS median_lcp,
  APPROX_QUANTILES(lcp, 100)[OFFSET(75)] AS p75_lcp,
  COUNT(*) AS sample_count
FROM
  `cloudflare-beacon.rum.web_vitals`
WHERE
  DATE(timestamp) >= DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY)
  AND lcp IS NOT NULL
GROUP BY
  country_code
HAVING
  sample_count > 1000
ORDER BY
  median_lcp ASC
LIMIT 50;

Analyzing Browser Performance

-- Compare LCP distributions across browser families
SELECT
  browser_family,
  APPROX_QUANTILES(lcp, 100)[OFFSET(50)] AS p50_lcp,
  APPROX_QUANTILES(lcp, 100)[OFFSET(75)] AS p75_lcp,
  APPROX_QUANTILES(lcp, 100)[OFFSET(95)] AS p95_lcp,
  COUNTIF(lcp <= 2500) / COUNT(*) AS good_lcp_rate,
  COUNT(*) AS total_measurements
FROM
  `cloudflare-beacon.rum.web_vitals`
WHERE
  DATE(timestamp) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
GROUP BY
  browser_family
ORDER BY
  total_measurements DESC;

The p75 metric is particularly important here — Google's CrUX (Chrome User Experience Report) uses the 75th percentile to determine whether a URL passes or fails Core Web Vitals thresholds. A "Good" LCP requires that at least 75% of your users experience LCP under 2.5 seconds, meaning your p75 must be below that threshold.


Translating Dataset Insights Into Site Improvements

Understanding global performance benchmarks is only useful if you can apply those learnings to your own properties. Here's a practical workflow:

Step 1: Establish Your Baseline

Before comparing yourself to BEACON benchmarks, you need your own performance data. Use synthetic testing to get reproducible measurements you can track over time. The Website Performance Analyzer provides Lighthouse-based audits that surface LCP, INP, CLS, and TTFB alongside actionable recommendations — a solid starting point before you invest in full RUM instrumentation.

Step 2: Implement RUM on Your Own Site

If you're not already collecting field data, start now. Options range from free (Google's web-vitals library sending to Analytics) to comprehensive (commercial RUM platforms):

// Minimal RUM setup using the web-vitals library
import { onLCP, onINP, onCLS, onFCP, onTTFB } from 'web-vitals';

function sendToAnalytics(metric) {
  const body = JSON.stringify({
    name: metric.name,
    value: metric.value,
    rating: metric.rating, // 'good', 'needs-improvement', or 'poor'
    delta: metric.delta,
    id: metric.id,
    navigationType: metric.navigationType,
  });
  
  // Use sendBeacon for reliability during page unload
  navigator.sendBeacon('/analytics', body);
}

onLCP(sendToAnalytics);
onINP(sendToAnalytics);
onCLS(sendToAnalytics);
onFCP(sendToAnalytics);
onTTFB(sendToAnalytics);

Step 3: Contextualize Against BEACON Benchmarks

Once you have your own p75 metrics, query the BEACON dataset to understand where you stand relative to sites in your category and serving similar geographies. If your p75 LCP is 3.1 seconds and the BEACON median for your target regions is 2.8 seconds, you're slightly below average — but you now know the gap is closeable.

Step 4: Prioritize High-Impact Fixes

Focus on the metrics with the largest gap from "Good" thresholds, weighted by traffic volume. A 500ms LCP improvement for users in your highest-traffic region will move your p75 more than a 2-second improvement for a small segment.

Common high-ROI optimizations:

  • Preconnect to critical origins: <link rel="preconnect" href="https://fonts.googleapis.com">
  • Self-host web fonts to eliminate third-party DNS resolution
  • Implement proper caching headers — the Cache Inspector can reveal if your static assets are being cached correctly
  • Use a CDN with edge nodes close to your user base
  • Defer non-critical JavaScript with async or defer attributes
  • Compress responses with Brotli (better than gzip for text assets)

Step 5: Monitor SSL and Security Fundamentals

Performance and security are intertwined. TLS handshake overhead contributes to TTFB, and misconfigured certificates can cause complete connection failures. Regularly verify your SSL configuration with the SSL Certificate Checker — an expired or misconfigured certificate doesn't just break security, it breaks performance for every affected user.


The SEO Connection: Why Performance Data Affects Rankings

Core Web Vitals aren't just a developer metric — they're a Google ranking signal. Page Experience signals, which include Core Web Vitals, mobile-friendliness, and HTTPS, factor into search rankings. The BEACON dataset's scale makes it a useful reference for understanding what Google sees across the web when evaluating these signals.

If you're working to improve rankings alongside performance, a comprehensive SEO Audit can surface issues beyond Core Web Vitals — meta tags, heading structure, canonical URLs, and structured data — that compound the performance work you're doing. Performance improvements alone won't move rankings if fundamental on-page SEO issues remain unaddressed.


Building a Performance Culture With Real Data

One of the less-discussed benefits of datasets like BEACON is their value in organizational conversations. Performance work often competes with feature development for engineering time, and abstract scores can be hard to translate into business impact.

Real user data changes that conversation. When you can show stakeholders that 40% of your users in a target market are experiencing "Poor" LCP — and that the BEACON dataset shows competitors in that region at "Good" — the business case for performance investment becomes concrete.

Establishing performance budgets tied to real user percentiles (rather than Lighthouse scores) also creates more meaningful accountability:

// performance-budget.json
{
  "metrics": {
    "lcp_p75": {
      "threshold": 2500,
      "unit": "ms",
      "source": "rum"
    },
    "inp_p75": {
      "threshold": 200,
      "unit": "ms",
      "source": "rum"
    },
    "cls_p75": {
      "threshold": 0.1,
      "unit": "score",
      "source": "rum"
    }
  }
}

Integrating these checks into CI/CD pipelines — failing builds that would regress field metrics — keeps performance from eroding incrementally over time.


Conclusion

The BEACON dataset is a significant contribution to the web performance community. Having access to billions of real user measurements across regions, browsers, and device types gives developers something that was previously only available to the largest tech companies: genuine population-level performance data to benchmark against.

But the dataset is a starting point, not an endpoint. The real work is instrumenting your own properties, understanding your specific user distribution, and systematically closing the gap between where you are and where "Good" looks like for your audience.

OpDeck's suite of tools can support that process at every stage — from initial audits with the Website Performance Analyzer to ongoing monitoring of SEO health with the SEO Audit and security verification with the SSL Certificate Checker. Start with a baseline audit at opdeck.co and see where your site stands against real-world performance standards.