Paper explainer · identification before optimisation论文解读 · 先识别,再优化

Can data from one observation protocol evaluate another that was never used?

一种观测协议采集的数据,能否评估另一种从未实施过的协议?

Not always. Even an infinite benchmark can be compatible with different latent temporal structures—and those structures can assign different predictive values to the same alternative protocol.

未必。即使基准数据趋于无穷,它仍可能同时兼容多种不同的潜在时间依赖结构;而这些结构会对同一个替代协议给出不同的预测价值。

A protocol's value is the fraction of target variation that the best possible predictor could explain from the measurements that protocol would collect.

协议价值,是理想预测器仅利用该协议所采集的测量时,能够解释的目标总体变异比例。

Counterfactual Evaluation of Temporal Observation Protocols
Xizhe Zhang张锡哲

Same benchmark同一个基准two latent worlds with identical measurement–target laws under protocol A两个潜在世界,在协议 A 下有完全相同的测量–目标分布 Two values两个价值I(B;K_+) = 0.682   I(B;K_−) = 0.827
Minimal mathematical example · Theorem 3最小数学例子 · 定理 3

Same benchmark. Two compatible worlds. Two different answers.

同一个基准。两个兼容世界。两个不同答案。

Protocol A observes one point in time. Two latent temporal models produce exactly the same distribution of A's measurement and the target. No amount of additional sampling under A can distinguish them.

协议 A 只观测一个时间点。两种潜在时间模型产生完全相同的“测量—目标”分布。继续在协议 A 下增加样本,仍无法区分它们。

Protocol B observes two other points. Its value is 0.682 in one compatible world and 0.827 in the other.

协议 B 改为观测另外两个时间点。在一个兼容世界中,它的价值是 0.682;在另一个世界中,是 0.827

The unresolved quantity is not how each new measurement relates to the target. It is how redundant the new measurements are with one another—dependence that protocol A never reveals.

未被确定的并不是每个新测量与目标之间的联系,而是这些新测量彼此之间有多冗余。协议 A 从未揭示这种依赖。

ε = 0last admissible ε (both worlds remain valid correlation matrices)可容许的上限(两个世界都仍是合法相关矩阵)

Benchmark quantities remain fixed基准数据能测到的量保持不变

K₊K₋
Var(Y_A)1.0000001.000000
Cov(Y_A, Θ)0.3882500.388250
Var(Θ)0.4280120.428012
max discrepancy最大差异 1e-16 · identical benchmark laws基准分布完全相同

Alternative-protocol value changes替代协议的价值在变

I(B; K_+)
0.682
I(B; K_−)
0.827
difference相差 0.1457

B observes Z_1 and Z_2; A observed only Z_0; the target is the mean of Z_0, …, Z_3.

B 观测 Z_1Z_2;A 只观测了 Z_0;目标是 Z_0, …, Z_3 的平均。

What moves什么在变

Only Cov(Z_1, Z_2) = ρ(1) ± ε changes along the invisible direction δ = (0, 1, −2, 1). Each measurement's link to the target is fixed; their overlap is not.

沿不可见方向 δ = (0, 1, −2, 1) 只有 Cov(Z_1, Z_2) = ρ(1) ± ε 在变。每个测量与目标的联系固定不变,变的是它们之间的重叠。

The benchmark determines the value of an undeployed protocol only when every latent model compatible with the benchmark assigns that protocol the same value.

只有当所有与基准数据相容的潜在模型都赋予替代协议同一个价值时,该价值才被基准数据识别。

Three questions, in the order they must be answered

三个问题,必须按顺序回答

Identify · Section 3识别 · 第 3 节

Is the value determined by the available data?

现有数据是否已经确定了这个价值?

Characterise the latent variation that the realised benchmark cannot see, and test whether that variation changes the value of the proposed protocol.

找出已实施基准看不见的潜在变化,并判断这些变化是否会改变拟议协议的价值。

Calibrate · Section 4校准 · 第 4 节

What additional measurements remove the relevant ambiguity?

补充哪些测量能够消除相关歧义?

Targeted augmentation can identify one specified value. A small densely observed calibration subset can support comparison across a broader family of protocols.

定向补测可以识别一个指定协议的价值;小规模密集校准子集则可以支持对更大协议族的比较。

Design · Section 5设计 · 第 5 节

Which protocol is best at the resolution the data can support?

在数据可支持的分辨率上,哪个协议更好?

Optimise timing, support, repetition, noise, and cost only after candidate values are identified and their differences are large enough to distinguish.

只有在候选价值可识别、且差异足以被区分之后,才优化测量时刻、窗口、重复、噪声与成本。

More subjects are not always the data you need.

真正缺少的,未必是更多受试者。

When an ambiguity is invisible under the realised protocol, increasing the number of units reproduces the same information. A smaller group observed more intensively may be more useful than many additional units observed in the same sparse way.

当某种歧义对已实施协议不可见时,继续增加样本只会重复同一组信息。与其在相同稀疏协议下继续扩样,一个规模较小但观测更密集的子集,可能更能支持未来协议的评估。

This suggests a practical study design: a large routine cohort paired with a smaller, densely observed calibration subset chosen with future protocol decisions in mind.

因此,一个可操作的研究设计是:大规模常规队列,配合一个专门为未来协议决策而设计的小规模密集校准子集。

Technical detail: how the calibration subset is used技术细节:校准子集怎么用

From m densely observed units, form the sample covariance, subtract the calibration noise, floor the eigenvalues so the result is a valid covariance, and rescale to a correlation matrix. Plugging the estimate into every candidate's value gives a family-uniform error of order C‖K̂ − K‖β (Theorem 10), falling at the root-m rate at a fixed model (Theorem 11).

m 个密集观测对象:算样本协方差、扣除校准噪声、给特征值设下限以保证合法协方差、再标准化为相关矩阵。把估计代入每个候选协议的价值,得到整个候选族上一致的误差 C‖K̂ − K‖β(定理 10),在固定模型下以根号 m 速率下降(定理 11)。

Large routine cohort大规模常规队列same sparse protocol, repeated across many units同一个稀疏协议,在很多对象上重复 Small dense calibration subset小规模密集校准子集reveals the cross-time dependence needed for future protocol comparisons揭示未来协议比较所需的跨时刻依赖
Corollary 12 · simulation under known data-generating laws推论 12 · 已知生成机制下的模拟

Calibration determines how finely protocols can be compared.

校准数据决定协议能够比较到多细。

A fine candidate family may contain better protocols, but its leading candidates can differ by very small amounts. Those differences are meaningful only when calibration error is smaller than the value gaps being compared.

更细的候选类可能包含更好的协议,但领先候选之间的价值差距也可能非常小。只有当校准误差小于这些差距时,这种精细比较才有意义。

If the largest protocol-value error is ε, the regret of selecting the empirical best is bounded by approximately , plus any optimisation error. Differences below that scale cannot be distinguished uniformly from calibration uncertainty.

若所有协议价值估计的最大误差为 ε,则按估计值选择协议的损失至多约为 ,再加上优化误差。低于这一尺度的差异,无法与校准不确定性稳定地区分。

The useful granularity of optimisation is set by the data, not by the optimiser.

优化能细化到什么程度,由数据决定,而不是由优化器决定。

True selection regret of the best estimated protocol in the coarsest class (2 layouts) and the finest class (568 exact supports), as calibration grows from 25 to 1000 units (Figure 2c). At m = 25 the finest class already wins ( vs ); at m = 1000 the coarsest is stuck at its restriction regret () while the finest reaches . 真实选择损失:最粗类(2 种布局)与最细类(568 个精确位置)中最优估计协议的损失,校准对象从 25 增至 1000(图 2c)。m = 25 时最细类已领先();m = 1000 时最粗类停在限制损失 ,最细类降到
Proposition 13 · Table 1 · simulation命题 13 · 表 1 · 模拟

The best observation schedule depends on the target.

最佳观测方案取决于要预测的目标。

A protocol that reconstructs the latent trajectory well need not be optimal for predicting a specific aggregate of that trajectory. The paper derives the exact marginal value of adding a measurement and uses it for cost-constrained, target-aware search.

能够较好重建整条潜在轨迹的协议,不一定最适合预测某个特定的轨迹聚合量。论文推导了新增一次测量所带来的精确边际价值,并据此进行预算约束下的目标特定搜索。

In 25 fully enumerated simulation settings, greedy search with one-swap refinement attained at least 98.7% of the target-specific optimum.

在 25 个可穷举的模拟设置中,贪心搜索加一次交换改进,至少达到目标特定最优值的 98.7%

Technical detail: the exact marginal gain技术细节:精确的边际收益

With residual covariance PS = K − QS, adding action a updates QS∪a = QS + vvᵀ/s, v = PSa, s = ℓaᵀPSa + ra; for the mean target the gain is (ωᵀPSa)²/s. The objective is monotone but not submodular, so the search is measured against exhaustive enumeration rather than given an approximation ratio.

记残差协方差 PS = K − QS,加入动作 aQS∪a = QS + vvᵀ/sv = PSas = ℓaᵀPSa + ra;均值目标下收益为 (ωᵀPSa)²/s。目标函数单调但不次模,因此用穷举对照来衡量搜索,而不声称近似比。

Change the target and watch the chosen times move — or drag any vertical line to another of the 16 candidate slots; the value is recomputed with the paper's formulas. 切换目标,看被选中的测量时刻如何变化;也可以把任意一条竖线拖到 16 个候选时刻中的另一个,价值会按论文公式实时重算。
Greedy target-aware selection of four point measurements from sixteen candidates on one latent process (recency-weighted target, noise 0.3). Numbers give the greedy order of selection; the faint row shows the other target's choice. The lines are draggable (they snap to the candidate slots), so you can test your own schedule against the greedy one. Computed in the browser with the paper's formulas. 同一个潜在过程上,从 16 个候选中贪心选出 4 个点测量(偏重近期的目标权重,噪声 0.3)。数字是贪心选中的顺序;淡色的一行是另一种目标的选择。竖线可以拖动(吸附到候选时刻),用来把你自己的方案和贪心方案比一比。在浏览器里用论文的公式现算。
Retrospective reconstruction from fully annotated data · Section 6.4基于完整标注数据的回溯重建 · 第 6.4 节

Coarse layout differences were clearer than exact learned locations.

粗粒度布局的差异,比精确学习的位置更清楚。

Sleep-EDF
0.682dispersed分散vs0.648contiguous · N = 4 epochs连续 · N = 4 个片段

At matched scoring budgets, dispersed epochs often outperformed a centred contiguous block in the pooled Sleep analyses. The magnitude and direction of the contrast varied across the two source studies.

在相同评分预算下,分散观测在合并的睡眠分析中常优于居中连续块,但差异的方向和幅度在两个来源研究之间存在异质性。

Long-Term AF
0.971dispersed分散vs0.696contiguous · four 15-min windows连续 · 四个 15 分钟窗口

Across the evaluated multi-window budgets, dispersed observation outperformed contiguous blocks. With four 15-minute windows, pooled held-out R² was 0.971 for dispersion and 0.696 for a contiguous block.

在所评估的多窗口预算中,分散观测均优于连续块。四个 15 分钟窗口时,分散方案的汇总留出 R² 为 0.971,连续块为 0.696

Learned supports学得的位置
−0.012median advantage over fixed dispersion · range [−0.126, +0.097]相对固定分散的优势中位数 · 区间 [−0.126, +0.097]

Exact target-aware Sleep supports were less stable across targets and subject resamples, and they did not show a stable held-out advantage over fixed dispersion.

精确的目标特定睡眠锚点在不同目标与受试者重抽样之间稳定性较弱,也没有显示出相对固定分散方案的稳定留出优势。

These analyses illustrate the paper's resolution result: broad protocol layouts can be distinguishable even when the data do not support reliable selection of exact anchor locations.

这些结果体现了论文的“分辨率”结论:数据可以支持对粗粒度协议布局的比较,却未必足以可靠选择精确锚点。

197 Sleep-EDF hypnograms (100 subjects) and 84 Long-Term AF records; expert annotation files only, subject-disjoint five-fold cross-fitting, pooled held-out R². Full charts in the technical tour.

197 段 Sleep-EDF 分期(100 名受试者)与 84 段长程房颤记录;只用专家标注文件,受试者不交叠的五折交叉拟合,汇总留出 R²。完整图表见技术导览

This is a general problem of changing what is observed.

这不仅是时间采样问题,而是“改变观测系统”时的普遍问题。

The same question arises whenever a proposed measurement has never been observed jointly with the target: longer or dispersed clinical monitoring, overlapping multimodal panels, wearable and imaging schedules, or new environmental sensor locations.

只要拟议测量从未与目标共同出现,同样的问题就会发生:更长或更分散的临床监测、重叠的多模态测量面板、可穿戴与影像采集方案,以及新的环境传感器位置。

In such settings, feature selection and sensor optimisation begin only after the counterfactual value of the proposed measurements has been identified or calibrated.

在这些场景中,特征选择与传感器优化都应当从一个更早的问题开始:这些拟议测量的反事实价值是否已经被现有数据识别或校准。

Clinical monitoring临床监测

Longer, repeated or dispersed recordings whose burden was never paired with the outcome.

更长、重复或分散的记录,其负担从未与结局一起被观测。

Multimodal panels多模态面板

Overlapping assay or biomarker panels that were collected on different subjects.

在不同对象上采集、彼此重叠的检测或生物标志物面板。

Wearables and imaging可穿戴与影像

Sampling schedules for devices and scans that decide how often and how long to observe.

决定设备与扫描多久看一次、看多长的采集方案。

Environmental sensing环境传感

New sensor locations or reading frequencies for exposure and exceedance-time targets.

为暴露量或超标时长目标新增的传感器位置或读数频率。

Identification before optimisation.

先识别,再优化。

Existing data do not automatically determine the value of measurements that were never made. The relevant goal is not always to recover the full latent system, but to remove the ambiguity that changes the protocol values under consideration. Calibration then determines which differences are resolvable, and design should proceed at that resolution.

现有数据不会自动确定那些从未实施过的测量的价值。真正需要恢复的未必是完整潜在系统,而是消除会改变候选协议价值的歧义。校准数据随后决定哪些差异可以被分辨,观测设计应当在这一分辨率上进行。