mantispy.tl.empirical_fdr

Contents

mantispy.tl.empirical_fdr#

mantispy.tl.empirical_fdr(adata, *, control_genes, score, group=None, alpha=0.01, criterion='p', key_added='empirical_fdr', copy=False)[source]#

Score each gene against a control set that should produce no phenotype.

Genes that are not expressed in the screened cell line cannot show a knockout phenotype, so the spread of their scores is an empirical null: a knockout that scores past what those genes reach is unlikely to be noise. This is the calibration PERISCOPE applies with non-expressed genes, generalized to any per-gene score and any control set (mantispy.io.unexpressed_genes() builds one).

The score is yours to choose and must rise with the strength of the phenotype. A count of the features that separate a gene from the controls works well, from a per-feature test such as the Mann-Whitney p-values of effect_size() or the moderated t of differential_features(). An unbounded count separates the controls from the rest; a bounded score such as the activity mean average precision of map() saturates, so a handful of off-target controls at its ceiling block the low tail. This calibrates against control genes that cannot respond, where hit_calling() calibrates each group against control wells; use this when a set of non-responding genes is the better null.

Parameters:
  • adata (AnnData) – One row per gene, with the score in obs and a column (or index) naming the gene.

  • control_genes (Collection[str]) – The genes that form the null, such as the unexpressed set. A bare string is rejected, since it would be read as a set of single characters.

  • score (str) – obs column holding the per-gene score, higher for a stronger phenotype.

  • group (str | None (default: None)) – obs column naming the gene, matched against control_genes; the index is used when None.

  • alpha (float (default: 0.01)) – The cutoff a gene must clear in criterion to be called a hit.

  • criterion (str (default: 'p')) – Which quantity is_hit reads, "p" or "q" (see Notes).

  • key_added (str (default: 'empirical_fdr')) – Prefix for the outputs.

  • copy (bool (default: False)) – Return a modified copy instead of mutating in place.

Return type:

AnnData | None

Returns:

None, or the modified copy. Writes uns["mantispy"][key_added] with group, score, is_control, pvalue, qvalue and is_hit, and joins obs[key_added + "_pvalue"], obs[key_added + "_qvalue"], obs[key_added + "_control"] and obs[key_added] back onto the rows. Control genes and genes with a missing score get a missing pvalue and qvalue and are never hits.

Raises:
  • TypeError – control_genes is a string rather than a collection of gene names.

  • ValueError – criterion is not one of CRITERIA, or no control gene is present in the object, or every control gene has a missing score.

  • KeyError – score, or group, is not an obs column.

Notes

Two quantities come out, answering different questions. The p-value is the share of control genes that reach a gene’s score or beyond, with one added to the count and to the control total so a gene past every control gets a small positive value rather than zero, so a cutoff on it fixes the rate at which a control gene is called, the way PERISCOPE sets its threshold. It does not account for the number of genes tested, so the share of real false positives in a p-value hit list is higher than the cutoff whenever many genes carry no phenotype. The q-value is a target-decoy false discovery rate: at a gene’s score it compares the fraction of controls reaching it with the fraction of all tested genes reaching it, and is made monotone from the permissive end. It estimates the share of the hit list that is false, so it is the honest rate to quote, and it can sit well above the p-value cutoff when the controls and the tested genes overlap. criterion picks which one is_hit reads; the other is written too.

The controls must behave like the tested genes under the null for either number to mean what it says. A few controls with a strong, reproducible score (off-target reagents, or genes the expression reference mislabels) sit in the extreme tail and raise the floor the q-value can reach; tightening the expression cutoff rarely removes them, since an off-target effect is not an expression artifact. With few controls the null is coarse and a warning fires.