mantispy.tl.empirical_fdr#
- mantispy.tl.empirical_fdr(adata, *, control_genes, score, group=None, alpha=0.01, criterion='p', key_added='empirical_fdr', copy=False)[source]#
Score each gene against a control set that should produce no phenotype.
Genes that are not expressed in the screened cell line cannot show a knockout phenotype, so the spread of their scores is an empirical null: a knockout that scores past what those genes reach is unlikely to be noise. This is the calibration PERISCOPE applies with non-expressed genes, generalized to any per-gene score and any control set (
mantispy.io.unexpressed_genes()builds one).The score is yours to choose and must rise with the strength of the phenotype. A count of the features that separate a gene from the controls works well, from a per-feature test such as the Mann-Whitney p-values of
effect_size()or the moderated t ofdifferential_features(). An unbounded count separates the controls from the rest; a bounded score such as the activity mean average precision ofmap()saturates, so a handful of off-target controls at its ceiling block the low tail. This calibrates against control genes that cannot respond, wherehit_calling()calibrates each group against control wells; use this when a set of non-responding genes is the better null.- Parameters:
adata (
AnnData) – One row per gene, with the score inobsand a column (or index) naming the gene.control_genes (
Collection[str]) – The genes that form the null, such as the unexpressed set. A bare string is rejected, since it would be read as a set of single characters.score (
str) –obscolumn holding the per-gene score, higher for a stronger phenotype.group (
str|None(default:None)) –obscolumn naming the gene, matched againstcontrol_genes; the index is used whenNone.alpha (
float(default:0.01)) – The cutoff a gene must clear incriterionto be called a hit.criterion (
str(default:'p')) – Which quantityis_hitreads,"p"or"q"(see Notes).key_added (
str(default:'empirical_fdr')) – Prefix for the outputs.copy (
bool(default:False)) – Return a modified copy instead of mutating in place.
- Return type:
- Returns:
None, or the modified copy. Writesuns["mantispy"][key_added]withgroup,score,is_control,pvalue,qvalueandis_hit, and joinsobs[key_added + "_pvalue"],obs[key_added + "_qvalue"],obs[key_added + "_control"]andobs[key_added]back onto the rows. Control genes and genes with a missing score get a missingpvalueandqvalueand are never hits.- Raises:
TypeError –
control_genesis a string rather than a collection of gene names.ValueError –
criterionis not one ofCRITERIA, or no control gene is present in the object, or every control gene has a missing score.KeyError –
score, orgroup, is not anobscolumn.
Notes
Two quantities come out, answering different questions. The p-value is the share of control genes that reach a gene’s score or beyond, with one added to the count and to the control total so a gene past every control gets a small positive value rather than zero, so a cutoff on it fixes the rate at which a control gene is called, the way PERISCOPE sets its threshold. It does not account for the number of genes tested, so the share of real false positives in a p-value hit list is higher than the cutoff whenever many genes carry no phenotype. The q-value is a target-decoy false discovery rate: at a gene’s score it compares the fraction of controls reaching it with the fraction of all tested genes reaching it, and is made monotone from the permissive end. It estimates the share of the hit list that is false, so it is the honest rate to quote, and it can sit well above the p-value cutoff when the controls and the tested genes overlap.
criterionpicks which oneis_hitreads; the other is written too.The controls must behave like the tested genes under the null for either number to mean what it says. A few controls with a strong, reproducible score (off-target reagents, or genes the expression reference mislabels) sit in the extreme tail and raise the floor the q-value can reach; tightening the expression cutoff rarely removes them, since an off-target effect is not an expression artifact. With few controls the null is coarse and a warning fires.