CoFi-E: Collision-Guided Fiber Refinement for Anytime-Valid Conformal Detection

Research Project 2026
Md Nazmul Kabir Sikder

Overview

Sequential conformal detectors combine model-agnostic nonconformity scores with finite-sample, anytime false-alarm control. Their power, however, passes through a single scalar score. If a distribution change leaves the score's distribution unchanged, no betting strategy on the resulting conformal ranks can detect it: the detector is perfectly valid and completely blind at the same time.

We call this failure a rank-fiber collision. A score collapses the observation space into fibers (groups of observations with the same score), and any change inside a fiber is discarded. The same failure arises in security monitors when malicious behavior preserves a deployed risk-score distribution while changing the richer telemetry underneath. A stronger betting function cannot fix this; the representation itself has to change.

Approach

CoFi-E (Collision-guided Fiber-refining Conformal E-detection) starts from an existing deployed score, or learns from scratch, and repairs it against a declared portfolio of alternative distributions:

  • Find hidden differences. Within each current score fiber, search for a witness split that separates clean data from a training alternative. The exact information recovered by a split depends only on four probabilities, so it can be estimated by counting.
  • Certify every split. Witness, selection, and audit data are disjoint. Independent audit data give an exact binomial lower bound on each proposed split's gain, and only a certified prefix of the refinement path is kept.
  • Never forget. Splits are permanent and never merged, so refinement cannot lose rank information the detector already exposed for any alternative.
  • Keep validity separate. The refined score and its bettor are frozen before fresh reference data are drawn, so learning changes power but never the conformal null law.

Theory

  • The information in a tie-smoothed conformal rank equals the information in the score's induced partition, and a Pearson projection splits total information into a visible part and an unresolved within-fiber collision.
  • With a weak-witness condition on the split dictionary, greedy refinement shrinks the total portfolio collision at a geometric rate; for finite alphabets the number of refinements needed is tight.
  • Transfer bounds quantify how much evidence carries over to an unseen anomaly family as a function of its distance from the convex hull of the training alternatives.

Evaluation

  • Controlled recovery: refinement recovers information the original score misses while keeping the null e-values and false-alarm rate at their nominal levels.
  • Collision attacks: implemented attacks preserve marginal, context-conditional, and joint score interfaces exactly, yet remain clearly detectable in the full observations.
  • Image and time-series benchmarks: on MVTec AD, UCR, and UEA datasets with held-out anomaly families, the audit certificates stay informative, but accuracy on unseen families trails standard one-class detectors. Including a family in the training portfolio measurably improves detection, matching what the transfer bounds predict.

The result is a certified way to repair detector blind spots, with family coverage setting the scope of its guarantees rather than a claim of universal anomaly-detection performance.