2026-08-12

Two-sample Phase 2 model

\[ X_{Ti}\overset{iid}{\sim}N(\mu_T,\sigma^2), \qquad X_{Cj}\overset{iid}{\sim}N(\mu_C,\sigma^2) \]

\[ \delta=\mu_T-\mu_C, \qquad Y=\bar X_T-\bar X_C \]

Let

\[ \kappa=\frac1{n_T}+\frac1{n_C}, \qquad \nu=n_T+n_C-2. \]

Then

\[ Y\sim N(\delta,\sigma^2\kappa), \qquad \frac{\nu S_p^2}{\sigma^2}\sim\chi^2_\nu, \]

with \(Y\perp S_p^2\).

Goal: estimate \((\delta,\sigma^2)\) after a positive Phase 2 selection event.

Positive selection and conditional likelihood

The program advances only when

\[ \boxed{Y>c}. \]

For an observed Go result \((y,s_p^2)\),

\[ L_c(\delta,\sigma) = \frac{f_Y(y\mid\delta,\sigma)f_{S_p^2}(s_p^2\mid\sigma)} {P_{\delta,\sigma}(Y>c)} \propto \frac{ \sigma^{-(\nu+1)} \exp\!\left[-\dfrac{\nu s_p^2+(y-\delta)^2/\kappa}{2\sigma^2}\right] }{\bar\Phi(a)}, \]

where \(a=(c-\delta)/(\sigma\sqrt\kappa)\). The MCLE maximizes \(L_c\) over \(\delta\in\mathbb R\) and \(\sigma>0\).

As \(y\downarrow c\), the MCLE satisfies

\[ \widehat a_{\rm MCLE} \sim \frac{s_p\sqrt\kappa}{y-c} \longrightarrow\infty, \qquad \bar\Phi(\widehat a_{\rm MCLE})\longrightarrow0. \]

The denominator vanishes near the Go boundary, so \(1/\bar\Phi(\widehat a_{\rm MCLE})\) can overpower the data-fitting terms.

Figure 1: a barely selected result produces a flat, remote maximum

Profile conditional log likelihood for a barely selected result

Paper Figure 1. Example 1 with \(c=0.33\) and \(y=0.331\).

As \(y\downarrow c\),

\[ \widehat\delta_{\rm MCLE} =c-\frac{s_p^2\kappa}{y-c} +o\{(y-c)^{-1}\} \longrightarrow-\infty. \]

Simultaneously,

\[ \widehat p_{\rm Go} =\bar\Phi(\widehat a) \longrightarrow0. \]

The fitted model explains an observed Go result using an effect under which Go would have been essentially impossible.

Figure 2: the pathology occurs in selected trials

Probability of an ill-posed MCLE

Paper Figure 2. Ill-posed means \(\widehat\delta_{\rm MCLE}<-10\).
  • With 10,000 selected replicates per true effect, the estimated probability is 4.33% at \(\delta=0\) and 1.29% at \(\delta=0.35\).

  • Kirby et al. report

\[ \widehat\delta_{\rm Kirby} =\max(0,\widetilde\delta_{\rm MCLE}). \]

  • Truncation prevents a reported negative estimate, but it does not repair the singular conditional-likelihood objective. With \(c>0\), the resulting point mass at zero also remains below the original Go threshold.

Temper only the selection adjustment

Full conditional log likelihood

\[ \ell_c(\delta,\sigma) =-(\nu+1)\log\sigma -\frac{\nu s_p^2+(y-\delta)^2/\kappa}{2\sigma^2} -\color{#B33A3A}{\log\bar\Phi(a)}. \]

Tempered conditional log likelihood

\[ \ell_\gamma(\delta,\sigma) =-(\nu+1)\log\sigma -\frac{\nu s_p^2+(y-\delta)^2/\kappa}{2\sigma^2} -\color{#006D77}{\gamma\log\bar\Phi(a)}, \qquad 0\leq\gamma<1. \]

  • \(\gamma=0\): unadjusted likelihood
  • \(0<\gamma<1\): partial selection adjustment
  • \(\gamma=1\): full conditional likelihood
The modification acts on the source of instability; it is not post-estimation truncation.

Main result: finite behavior at the Go boundary

Let \(\lambda(a)=\phi(a)/\bar\Phi(a)\) and \(m_\gamma(a)=\gamma\lambda(a)-a\). For every fixed \(0\leq\gamma<1\), \(m_\gamma\) has a unique finite zero \(a_\gamma\):

\[ \boxed{a_\gamma=\gamma\lambda(a_\gamma)}. \]

As \(y\downarrow c\),

\[ \widehat a\longrightarrow a_\gamma, \qquad \widehat\sigma_\gamma^2 \longrightarrow\frac{\nu}{\nu+1}s_p^2, \]

and

\[ \boxed{ \widehat\delta_\gamma \longrightarrow c-a_\gamma s_p\sqrt{\frac{\kappa\nu}{\nu+1}} }. \]

Any fixed amount of tempering keeps the estimate finite. At \(\gamma=1\), no finite root exists and the full-MCLE singularity returns.

Why tempering works: a Gaussian factor remains

At \(y=c\) and fixed \(\sigma\), the \(a\)-dependent likelihood factor is

\[ \frac{\phi(a)}{\bar\Phi(a)^\gamma}. \]

Using \(\bar\Phi(a)\sim\phi(a)/a\),

\[ \frac{\phi(a)}{\bar\Phi(a)^\gamma} \sim \underbrace{a^\gamma}_{\text{polynomial}} \underbrace{\phi(a)^{1-\gamma}}_{\text{remaining Gaussian factor}}. \]

Tempered: \(\gamma<1\)

\(\phi(a)^{1-\gamma}\) drives the objective to zero as \(a\to\infty\) and prevents escape to the boundary.

Full MCLE: \(\gamma=1\)

The Gaussian factor is cancelled exactly; the objective can support \(a\to\infty\) and \(\widehat\delta\to-\infty\).

Far above the threshold, the correction disappears

The tempered score equation gives

\[ \widehat\delta_\gamma =y-\gamma\widehat\sigma_\gamma\sqrt\kappa\, \lambda(\widehat a). \]

When the fitted effect is far above \(c\),

\[ \widehat a\to-\infty, \qquad \lambda(\widehat a)\to0, \]

so

\[ \boxed{\widehat\delta_\gamma-y\longrightarrow0}. \]

If \(Y\overset{p}{\to}\delta\), then \(\widehat\delta_\gamma\overset{p}{\to}\delta\).

Tempering is active where selection matters and asymptotically negligible where it does not.

Select \(\gamma\) through an interpretable probability floor

Instead of choosing \(\gamma\) directly, specify

\[ \boxed{ \bar\Phi(\widehat a) =P_{\widehat\delta,\widehat\sigma}(Y>c) \geq p_{\min} }. \]

\(p_{\min}\) is the smallest pre-selection Go probability that the fitted model is allowed to assign to a selected trial. It is a property of the estimator—not a p-value, a posterior probability, or a claim about the true Go probability.

Choose

\[ a_\gamma=\Phi^{-1}(1-p_{\min}), \qquad \boxed{ \gamma=\frac{a_\gamma}{\lambda(a_\gamma)} =\frac{a_\gamma p_{\min}}{\phi(a_\gamma)} }. \]

For Example 1, \(p_{\min}=0.05\) gives \(\gamma=0.797\) and a boundary-limit estimate of approximately \(0.003\) when \(s_p=1\).

Simulation—Example 1: tempering strength

\(n_T=n_C=50\), \(\sigma=1\), \(c=0.33\); 10,000 selected replicates at each \(\delta\in\{0,0.05,\ldots,1\}\).

Mean bias for tempered MCLE in Example 1

Paper Figure 3. Conditional mean bias.

Adjusted estimates below threshold in Example 1

Paper Figure 5. Adjusted estimate below \(c\).
  • Every tempered estimator has finite mean bias.
  • Smaller \(p_{\min}\) means larger \(\gamma\) and stronger correction: less positive bias for weak effects, but more estimates below \(c\).
  • At \(\delta=0.35\), mean bias is \(-0.020\), \(0.017\), and \(0.040\) for \(p_{\min}=0.01\), \(0.05\), and \(0.10\), respectively.

Simulation—Example 1: Kirby comparison

Mean bias comparison with Kirby corrections

Paper Figure 9. Conditional mean bias.
  • At \(\delta=0.35\), tempered MCLE has mean bias \(0.017\), compared with \(0.042\) for LWB, \(0.077\) for CLWB, \(-0.023\) for percentile bias, and \(-0.063\) for Kirby MCLE.

  • Kirby MCLE corrects most strongly, but truncation at zero contributes to negative bias over much of the positive-effect range.

  • Tempered MCLE remains smooth and untruncated, with bias close to zero around and above \(c\).

No method is uniformly pointwise dominant.

Why compare two Phase 2/Phase 3 strategies?

Both strategies start with the same original Phase 2 Go population, \(Y>c\), and use the adjusted estimate \(d\) for Phase 3 planning.

Original Go: \(Y>c\) → adjusted estimate \(d\) → two possible actions

Strategy A: withdraw Go

  • Cancel Phase 3 when \(d\leq c\).
  • A cancelled program contributes zero achieved power.
  • Otherwise, size Phase 3 using \(d\).

Strategy B: retain Go

  • Keep the original Go decision regardless of \(d\).
  • Use \(d\) for Phase 3 sizing.
  • Assess economic feasibility separately.
Simulation question: does revising the original Go decision improve or reduce average achieved power across the original \(Y>c\) population?

Simulation—Example 1: preserve the Go decision

Strategy B: retain the original \(Y>c\) decision; use the adjusted estimate for Phase 3 planning. At \(\delta=0.35\), achieved power is 71.1% versus 28.0% under Strategy A for tempered MCLE.

A very small adjusted effect may imply an economically infeasible trial—this does not force launch. Use feasibility review, a prespecified minimum clinically/market-relevant design effect, or Phase 3 sample-size reassessment. Strategy A instead risks discarding a genuinely large effect after an extreme downward adjustment.

Takeaways

  1. Full MCLE can explain a barely positive result through a fitted Go probability approaching zero, producing an unbounded negative estimate.

  2. Replacing \(-\log\bar\Phi(a)\) by \(-\gamma\log\bar\Phi(a)\) with \(\gamma<1\) removes the boundary singularity while preserving recovery of the unadjusted estimate far above \(c\).

  3. Calibrating \(\gamma\) through \(p_{\min}\) makes the regularization interpretable and prespecifiable.

  4. For development decisions, let observed \(Y\) determine Phase 2 Go/No-go; use the adjusted estimate for Phase 3 planning, with feasibility safeguards when the adjusted design effect is small.

Stable correction at the boundary; negligible correction when selection is negligible.