mantispy.tl.aggregate_guides#
- mantispy.tl.aggregate_guides(adata, *, score, guide, gene, control, method='stouffer', weight=None, direction=None, n_null=10000, alpha=0.05, key_added='gene_aggregation', seed=0, copy=False)[source]#
Combine a gene’s guides into one hit call, calibrated against non-targeting controls.
A gene in a pooled CRISPR screen is targeted by many guides, and a call must pool them rather than treat each guide as its own result. This takes a one-sided per-guide p-value, small when a guide shows a phenotype, combines the guides of each gene into one statistic, and calibrates that statistic against random same-size groups of non-targeting-control (NTC) guides. Matching the group size matters because the combined statistic’s null depends on how many guides a gene has.
guide_activity()produces the per-guide p-value this reads, scoring each guide against a reference control class. The scoring reference and thecontrolnull must be two different control classes, or the null is scored against itself; a screen that carries both intergenic and non-targeting guides can set the scale with one and keep the other as the null."stouffer"turns each p into a z and sums the z’s, so it rewards a consistent effect across a gene’s guides;"fisher"combines the p’s and fires when any one guide is strongly significant, which is more powerful for a gene with a single potent guide but lets one off-target reagent carry the gene to a false call. This combines per-guide evidence within a gene, whereempirical_fdr()calibrates an already-per-gene score against a set of non-responding genes.- Parameters:
adata (
AnnData) – One row per guide, with the score, the guide id and the gene inobs.score (
str) –obscolumn holding the one-sided per-guide p-value in[0, 1], small for a stronger phenotype.guide (
str) –obscolumn naming the guide; one scored row per guide is expected, and a repeated id warns.gene (
str) –obscolumn naming the gene the guide targets.control (
str) – The value ofgenethat marks the non-targeting guides forming the null, such as"nontargeting".method (
str(default:'stouffer')) –"stouffer"(default) or"fisher"(see above and Notes).weight (
str|None(default:None)) –obscolumn of per-guide weights for a weighted Stouffer combination, such as a replicate or cell count; equal weights whenNone. Only"stouffer"accepts weights.direction (
str|None(default:None)) –obscolumn whose sign says which way a guide moved, so guides that disagree cancel instead of adding. Only"stouffer"uses it; a one-sided p already fixes the direction for"fisher".n_null (
int(default:10000)) – Random NTC groups drawn per distinct guide count to build the null.alpha (
float(default:0.05)) – The q-value cutoff a gene must clear to be called a hit.key_added (
str(default:'gene_aggregation')) – Prefix for the outputs.seed (
int(default:0)) – Seed for the null sampling.copy (
bool(default:False)) – Return a modified copy instead of mutating in place.
- Return type:
- Returns:
None, or the modified copy. Writesuns["mantispy"][key_added]with one row per gene:gene,n_guides,statistic,pvalue,qvalueandis_hit, sorted bypvalue. Joinsobs[key_added + "_pvalue"],obs[key_added + "_qvalue"]andobs[key_added]onto the guide rows, broadcast from each guide’s gene. Non-targeting guides, guides with a missing score, and genes left with no scored guide get a missingpvalueandqvalueand are never hits.- Raises:
ValueError –
methodis not one ofMETHODS,weightordirectionis given with"fisher",alphais not in(0, 1),n_nullis below one, a finitescorefalls outside[0, 1](so it is not a p-value), no control guide is present, or no gene has a scored guide.KeyError –
score,guide,gene,weightordirectionis not anobscolumn.
Notes
The p-value is the share of same-size NTC groups whose combined statistic reaches the gene’s or beyond, with one added to the count and the total so a gene past every NTC group gets a small positive value rather than zero. Each NTC group is drawn without replacement, a genuine same-size sample of the controls, so the pool of non-targeting guides must be well above the largest gene’s guide count; a warning fires when a gene uses more than half the pool, since a same-size group is then nearly the whole pool, its null has almost no spread, and its p-value cannot be trusted (a gene with more guides than the whole pool falls back to sampling with replacement). The q-value is the Benjamini-Hochberg false discovery rate over the tested genes, made monotone from the permissive end and never below the gene’s own p-value.
The non-targeting guides must behave like a true-null gene for the calibration to hold. With few NTC guides the null is coarse and a warning fires. A gene whose guides disagree in direction scores low under
"stouffer"unlessdirectionlets them cancel; under"fisher"a one-sided p already points one way, so opposing guides cannot cancel.