# The Manual-Sync Tax benchmark methodology

Status: complete controlled benchmark.

Protocol version: `2026-08-28.2`

The executable protocol was frozen before the full run. SHA-256 hashes are recorded in `research/manual-sync-tax/protocol-manifest.json`. This is a local integrity record, not third-party preregistration.

Sponsor and operator: Watch Peak Party. No external funding was used. This is original, reproducible WatchPeak research, not independent research.

## Research question

Under frozen browser, network-delay and buffering conditions, how far apart do two browser-video players become after a manual countdown compared with a projected synchronization method?

## Apparatus

Two isolated Playwright Chromium browser contexts load the same locally generated 20-second WebM clip. The clip is served from a temporary local HTTP server. It contains synthetic video and audio and uses no copyrighted programme content, streaming credentials, customer data or real watch-party sessions.

## Conditions

- Methods: manual countdown and WatchPeak projected synchronization.
- Simulated one-way latency: 0, 50, 150 and 300 milliseconds.
- Jitter: low (plus or minus 8 milliseconds) and high (plus or minus 65 milliseconds).
- Buffer pattern: no forced stall or a 900-millisecond follower stall.
- Sampling: every 100 milliseconds.
- Pilot: three runs per condition.
- Frozen full study: 30 runs per condition.

All pseudo-random timing values are seeded from the protocol version, profile, condition and repeat. Re-running the same protocol produces the same planned delays.

## Manual condition

After a common countdown target, each browser receives an independently seeded reaction delay between 70 and 310 milliseconds. Both play from time zero. No later correction is sent.

## WatchPeak condition

The host starts at the countdown target. A command reaches the follower after the frozen one-way delay and directs it to the projected current host position. Follow-up snapshots are sent 650 and 1,800 milliseconds after the play event. The playback authority sends another snapshot every 10 seconds. After a forced stall, the follower uses the production request retry ladder, subject to its four-second recent-snapshot guard. A correction is applied when absolute drift exceeds 550 milliseconds.

The timing constants are pinned to `extension/content.js` by an automated regression test. The harness exercises the production-derived timing policy against neutral browser video; it does not load streaming-service adapters or claim that every service behaves exactly like the local player.

## Measures

- Median absolute playback drift in milliseconds.
- 95th-percentile absolute playback drift.
- Maximum absolute playback drift.
- Marker-crossing difference, used as the reaction-spoiler proxy.
- Number of synchronization corrections.
- Recovery time after the forced follower stall.
- Failure status and reason.

At the method level, the reported 95th-percentile figure is the median of each successful run's P95 drift. Grouped values use the same aggregation rule.

The published `full-summary.json` names each aggregate explicitly:

| Field | Meaning |
|---|---|
| `medianDriftMs` | Median across runs of each run's median drift |
| `p95DriftMs` | Median across runs of each run's own P95 drift, not a pooled P95 |
| `medianRunMaxDriftMs` | Median across runs of each run's maximum sampled drift |
| `maxDriftMs` | The single largest drift sample in that slice |

### Correction to the frozen reporting code

The frozen benchmark code aggregates each group with `quantile(..., 0.5)` and published the result of that under the field name `maxDriftMs`, which made it a median of per-run maxima rather than a maximum. The consequence was visible: the method-level "max" for the WatchPeak policy was 524 ms while several of its own groups reported higher "maxes".

Because `research/manual-sync-tax/benchmark-core.mjs` is hashed in the pre-run integrity manifest, it was not edited after data collection. The published grouped summary is instead rebuilt from the unmodified row-level `runs.json` by `scripts/build-research-summary.mjs`, which emits the value under its correct name `medianRunMaxDriftMs` and adds a true `maxDriftMs`. No measurement, no row and no hash changed. The archived frozen `summary.json` in the results directory retains the original field name.

## Results

All 960 planned runs completed successfully: 480 manual-countdown runs and 480 WatchPeak-policy runs.

The design allocates exactly half the runs to a forced follower stall, and the two buffer conditions produce very different distributions. Results are therefore reported per condition rather than pooled:

| Buffer condition | Manual median | WatchPeak median | Reduction |
|---|---|---|---|
| No forced stall (480 runs) | 69.4 ms | 2.8 ms | 96.0% |
| 900 ms follower stall (480 runs) | 889.8 ms | 61.5 ms | 93.1% |

Pooling all 960 runs gives 440.4 milliseconds for manual countdown and 34.9 milliseconds for the WatchPeak policy, a 92.1% reduction, with median run-level P95 falling from 446.1 to 110.1 milliseconds. That pooled median is reported for completeness but is not used as the headline: it is the midpoint of the two clusters above, an artefact of the 50/50 design split, and it describes neither condition.

The tail did not disappear. The largest sampled drift in any successful run was 1,092 milliseconds for manual countdown and 913.3 milliseconds for the WatchPeak policy. Under the forced-stall condition, the WatchPeak median run-level P95 was 490.3 milliseconds. These values are part of the result, not exclusions.

## Exclusions

Runs are not silently deleted. A run is marked failed when one or both browser players do not make sufficient progress. Failed runs remain in the CSV and JSON output and are excluded only from timing aggregates.

## Claim rules

- Full results describe this controlled setup, not every device, network, title or streaming platform.
- The benchmark does not measure emotional closeness, relationship quality, loneliness or health outcomes.
- WatchPeak funding and operation must be disclosed whenever the findings are cited.
- Material null, adverse and failed results remain in the published package.

## Reproduction

From the project root:

```bash
npm install
npx playwright install chromium
npm run research:smoke
npm run research:pilot
npm run research:full
```

`ffmpeg` must be available on the local `PATH`. The runner uses it to create the neutral WebM fixture in a local cache.
