Back to Blog
Best Practices December 1, 2025 4 min read

Plan Your January Tests Now: How to A/B Test Header Bidding Changes Properly

Post-holiday January is the best time to test wrapper changes, but only if the test is designed well. Here is how to run header bidding experiments that produce answers you can trust.

HR
HBDR Research
December 1, 2025

Why plan now for January

Cyber Monday is today, and most ad ops teams are still in holiday freeze. That makes this the right moment to plan the tests you will run once the freeze lifts. January is usually the best testing window of the year: advertiser demand typically drops after the holidays, which makes each test less expensive, and traffic is steady enough to reach conclusions within a few weeks. Teams that arrive in January with a written test plan get a quarter of learning. Teams that start planning in January often lose most of it.

For B2B publishers the case is even stronger. Their audiences return to work in early January, and many business advertisers start new annual budgets in the first quarter. Getting the stack tuned before that demand ramps up pays off for the rest of the year.

What is worth testing

Good header bidding tests change one meaningful variable at a time. Common candidates include:

  • Auction timeouts. The Prebid timeout documentation explicitly recommends A/B testing to find the balance between letting bids arrive and keeping pages fast.
  • Bidder lineup. Removing a low-value partner, or moving one from client-side to server-side.
  • Floor models. Different floor levels or rule structures by segment.
  • Lazy loading margins and refresh policies. How early slots are requested, and when they refresh.
  • Ad layout. Unit count, spacing and positions on key templates.

Designing a test you can trust

Randomize by user, not by pageview

Assign each user to a group and keep them there for the length of the test, usually with a first-party cookie or a stable ID. If users switch groups between pageviews, effects on session behavior, frequency capping and bidder learning get mixed together.

Pick one primary metric

The best primary metric for most tests is revenue per thousand sessions, including all demand sources, header bidding, ad server demand and direct. CPM on its own is misleading: floor increases raise CPM while lowering fill, and more ad units can raise revenue per page while lowering pages per session. Use CPM, fill rate, timeout rate, viewability and page performance as secondary diagnostics.

Include the ad server in the result

A wrapper change that increases header bidding revenue can reduce ad server revenue from other channels by the same amount. Pass a key-value with the test group to the ad server so you can report total revenue per group, not just Prebid wins.

Run for full weeks

Traffic and demand differ by day of week. Run tests for at least two full weeks, and avoid windows with known anomalies, such as the first days of January when demand resets, or major news events that distort traffic.

Size the groups sensibly

Small effects need large samples. If you expect a change of a few percent in revenue per session, a 10% test group on a small site may never produce a clear result. Estimate the effect size you care about and check that your traffic can detect it before starting.

Write down the decision rule first

Before launching, decide what result would lead you to adopt the change, what would lead you to reject it, and what you will do if the result is inconclusive. It prevents the common trap of rationalizing whatever the numbers show.

Built-in tools that help

Prebid includes some testing support. The Price Floors module offers skipRate, which skips floors on a random share of auctions, and modelGroups with weights, which lets you run multiple floor models at once and log which model applied. For other changes, the usual approach is a small bucketing script that sets the test group before the wrapper loads, then configures Prebid and the ad server key-value accordingly.

Report results so others can use them

Keep a short record for every test: the hypothesis, the groups and their sizes, the dates, the primary and secondary results, and the decision. Over a year, that log becomes the most valuable reference your team has, because it explains why each setting is what it is and stops the same test from being run twice.

Common mistakes

  • Testing several changes at once. If group B has a new timeout and a new bidder, you cannot tell which one helped or hurt.
  • Stopping early. Early results swing widely. Stopping as soon as a test looks good inflates the apparent win.
  • Ignoring page performance. A change that raises revenue but slows pages may cost more in lost sessions over time than it gains. Include Core Web Vitals in your secondary metrics.
  • Forgetting seasonality. A test that starts in the last week of December and ends in mid-January is measuring the holiday drop as much as the change.
  • Leaving tests running. Once a decision is made, roll the winner out to everyone and remove the test code.

A simple January plan

  1. Week 1: Finish post-holiday reviews, reset holiday floors, and confirm baselines.
  2. Weeks 2–3: Run the highest-priority test, usually timeouts or bidder lineup.
  3. Weeks 4–5: Run the second test, often floor structure for the new year.
  4. End of month: Document results, roll out winners, and schedule the next round.

The takeaway

Most wrapper settings in production were chosen once and never tested. January is the cheapest time of year to find out whether they are right, as long as the tests are designed well: one variable, user-level randomization, total revenue per session as the primary metric, and enough time and traffic to reach a clear answer. HBDR runs structured tests like these with the publishers it works with, and early December is when that planning starts.

Tags: a/b testing prebid header bidding analytics b2b publishers

Ready to maximize your ad revenue?

Get Started