GeneBench‑Pro presents ten case studies that illustrate representative benchmark tasks, providing the original prompts, the datasets used, and supporting materials. File previews show excerpts from the full datasets.
What the case studies contain
Each of the ten tasks covers a distinct laboratory or clinical genomics problem. Every case study includes the original problem statement, the data made available for the task, and the supplementary files. The materials note that several tasks use synthetic labels (for example TXR1, LINC473, KIN1, PROTA, PROTB, DRX1, Locus Q) and that any resemblance to real human genes is coincidental.
Highlighted examples among the case studies
-
Estimating clinical utility for a TXR1‑directed inhibitor: using a molecular tumor board registry, participants must estimate the marginal effect of TXR1i versus non‑TXR1 systemic therapy on week‑16 clinical benefit for tumors with SV‑driven TXR1 activation at baseline, and the 8‑week treatment‑limiting toxicity/discontinuation risk under TXR1i. The target subgroup must be reconstructed from long‑read, expression, tumor‑quality and pharmacogenomic evidence. A net clinical utility metric is defined as benefit risk difference (percentage points) minus 0.35 times the toxicity risk (percentage points), and a binary therapy_class_code is required.
-
Determining whether an apparent LINC473 lncRNA dependency is transcript‑specific: pooled CRISPRi screens, guide‑level local expression, CasRx transcript‑targeting follow‑ups, and single‑guide growth assays are provided. The analysis must control for local DNA‑locus perturbation, neighbor‑gene repression, guide swaps, GC toxicity and plate effects.
-
Estimating direct disease effects for nearby proteins with cis‑MVMR: two correlated proteins (PROTA and PROTB) share a locus; the task is to move from marginal associations to LD‑aware conditional log‑odds effects per +1 SD increase in log10 protein concentration for each protein.
-
Ancestry‑specific carrier frequencies and residual reproductive risk for DRX1: using screening and calibration files, pseudogene‑aware calls and ancestry‑specific assay calibration, participants must report AFR and EUR carrier frequencies among screening roster adults, the residual carrier risk for an AFR individual with a negative screen, the partner carrier frequency for a uniformly sampled partner, and the couple reproductive risk when the index is AFR and screen‑negative.
-
Genotype effect on activated‑monocyte CXCL10 expression from scRNA‑seq: because ambient RNA contaminates both target expression and activation markers, decontamination must precede the eQTL model. The goal is the per‑allele log rate ratio for CXCL10 in the activated monocyte subpopulation.
-
Calibrated association for a nested structural subhaplotype in Locus Q: separate the nested segment‑B calibrated copy‑dosage signal from the broader inversion orientation; report the cohort‑level log OR per segment‑B copy, the expression log fold change for the supported gene, a support code (1 if expression_log_fc > 0 and the clinical association is protective), and the number of calibrated carriers.
-
Hi‑C loop strength contrast after artifact masking: using 20 kb and 40 kb contact matrices and bin annotations, estimate loop enrichment at the 20 kb interaction between bin_id = 8 and bin_id = 17, reporting mean log2(observed/expected) across case replicates, across control replicates, and their difference.
-
Mapping a chromosome‑1 QTL in an eight‑founder recombinant population: reconstruct founder ancestry from biallelic markers, check marker orientations and separate the QTL from a batch‑aligned nuisance peak; report position (cM) and which founder carries the high‑effect allele.
-
Parent‑specific ancestry proportions and admixture timing from phased local‑ancestry tracts: after repairing reciprocal tract artifacts and a chromosome‑specific label inversion, estimate for each transmitted parental haplotype the fraction of ancestry A and the generations since a single pulse admixture; label parent1 as the haplotype with the smaller ancestry‑A fraction.
-
Inferring which of two haploid loci is under stronger positive selection from ancient allele‑frequency time series: place both loci on the same derived‑allele scale, model sample‑level sequencing error (~1%), and estimate the selection coefficient s (> 0) for the more strongly selected locus (selected_locus = "A" or "B").
Common requirements and scoring
Many tasks require not only numerically correct outputs but also clear, reproducible analytical reasoning: justification of preprocessing choices, handling of confounders, and acknowledgement of limitations. Several case studies explicitly demand that the final answer be returned as a single JSON object with specified keys and no additional prose.
Why this matters
The case studies span the interface between laboratory genomics and clinical decision support. By providing reproducible problem instances and supporting data, GeneBench‑Pro aims to enable method comparisons, reproducibility checks and evaluation of the quality of analytical reasoning in addition to numerical accuracy.
Note on synthetic labels and evaluation
The documentation reiterates that some labels are synthetic and that the provided datasets originate from real experiments. Submissions are evaluated not only for final numbers but for the strength and clarity of the analytical approach used to derive them.



