mantispy.tl.aggregate_guides

mantispy.tl.aggregate_guides#

mantispy.tl.aggregate_guides(adata, *, score, guide, gene, control, method='stouffer', weight=None, direction=None, n_null=10000, alpha=0.05, key_added='gene_aggregation', seed=0, copy=False)[source]#

Combine a gene’s guides into one hit call, calibrated against non-targeting controls.

A gene in a pooled CRISPR screen is targeted by many guides, and a call must pool them rather than treat each guide as its own result. This takes a one-sided per-guide p-value, small when a guide shows a phenotype, combines the guides of each gene into one statistic, and calibrates that statistic against random same-size groups of non-targeting-control (NTC) guides. Matching the group size matters because the combined statistic’s null depends on how many guides a gene has.

guide_activity() produces the per-guide p-value this reads, scoring each guide against a reference control class. The scoring reference and the control null must be two different control classes, or the null is scored against itself; a screen that carries both intergenic and non-targeting guides can set the scale with one and keep the other as the null. "stouffer" turns each p into a z and sums the z’s, so it rewards a consistent effect across a gene’s guides; "fisher" combines the p’s and fires when any one guide is strongly significant, which is more powerful for a gene with a single potent guide but lets one off-target reagent carry the gene to a false call. This combines per-guide evidence within a gene, where empirical_fdr() calibrates an already-per-gene score against a set of non-responding genes.

Parameters:
  • adata (AnnData) – One row per guide, with the score, the guide id and the gene in obs.

  • score (str) – obs column holding the one-sided per-guide p-value in [0, 1], small for a stronger phenotype.

  • guide (str) – obs column naming the guide; one scored row per guide is expected, and a repeated id warns.

  • gene (str) – obs column naming the gene the guide targets.

  • control (str) – The value of gene that marks the non-targeting guides forming the null, such as "nontargeting".

  • method (str (default: 'stouffer')) – "stouffer" (default) or "fisher" (see above and Notes).

  • weight (str | None (default: None)) – obs column of per-guide weights for a weighted Stouffer combination, such as a replicate or cell count; equal weights when None. Only "stouffer" accepts weights.

  • direction (str | None (default: None)) – obs column whose sign says which way a guide moved, so guides that disagree cancel instead of adding. Only "stouffer" uses it; a one-sided p already fixes the direction for "fisher".

  • n_null (int (default: 10000)) – Random NTC groups drawn per distinct guide count to build the null.

  • alpha (float (default: 0.05)) – The q-value cutoff a gene must clear to be called a hit.

  • key_added (str (default: 'gene_aggregation')) – Prefix for the outputs.

  • seed (int (default: 0)) – Seed for the null sampling.

  • copy (bool (default: False)) – Return a modified copy instead of mutating in place.

Return type:

AnnData | None

Returns:

None, or the modified copy. Writes uns["mantispy"][key_added] with one row per gene: gene, n_guides, statistic, pvalue, qvalue and is_hit, sorted by pvalue. Joins obs[key_added + "_pvalue"], obs[key_added + "_qvalue"] and obs[key_added] onto the guide rows, broadcast from each guide’s gene. Non-targeting guides, guides with a missing score, and genes left with no scored guide get a missing pvalue and qvalue and are never hits.

Raises:
  • ValueError – method is not one of METHODS, weight or direction is given with "fisher", alpha is not in (0, 1), n_null is below one, a finite score falls outside [0, 1] (so it is not a p-value), no control guide is present, or no gene has a scored guide.

  • KeyError – score, guide, gene, weight or direction is not an obs column.

Notes

The p-value is the share of same-size NTC groups whose combined statistic reaches the gene’s or beyond, with one added to the count and the total so a gene past every NTC group gets a small positive value rather than zero. Each NTC group is drawn without replacement, a genuine same-size sample of the controls, so the pool of non-targeting guides must be well above the largest gene’s guide count; a warning fires when a gene uses more than half the pool, since a same-size group is then nearly the whole pool, its null has almost no spread, and its p-value cannot be trusted (a gene with more guides than the whole pool falls back to sampling with replacement). The q-value is the Benjamini-Hochberg false discovery rate over the tested genes, made monotone from the permissive end and never below the gene’s own p-value.

The non-targeting guides must behave like a true-null gene for the calibration to hold. With few NTC guides the null is coarse and a warning fires. A gene whose guides disagree in direction scores low under "stouffer" unless direction lets them cancel; under "fisher" a one-sided p already points one way, so opposing guides cannot cancel.