Algorithms and theory
Empirical scaling of a solver on an instance family
Where a problem family becomes intractable on this machine, with fitted scaling across solvers and the theory that explains it.
Algorithms and theory
Graph-algorithm study on generated families
Runtime and solution quality of graph algorithms across generated graph families and sizes.
Algorithms and theory
SMT solver benchmarking with validated answers
Solver comparison on SMT-LIB with every model and unsat core checked.
Algorithms and theory
Symbolic computation behind a proof
Closed forms, recurrences and inequalities checked symbolically to support an argument.
Algorithms and theory
Anytime evaluation of a heuristic optimiser
Quality-vs-time curves of heuristics on TSPLIB-style instances against known optima.
Compilers and languages
Is this peephole rewrite sound?
An SMT proof or counterexample that an optimisation preserves semantics for all inputs.
Compilers and languages
Parser, grammar and AST work
Grammar written and tested, or ASTs of real code analysed.
Compilers and languages
LLVM IR and optimisation-pipeline study
IR built and run through optimisation passes, with before/after IR and JIT results.
Computer architecture
Roofline model of a kernel
Arithmetic intensity vs attainable performance for a kernel on a named CPU/GPU from its datasheet numbers.
Computer vision
Score detections against ground truth
COCO-style mAP/AR for a set of detections the researcher has.
Computer vision
Classical image processing and matching
Feature detection, matching, homography and geometric correction on images.
Computer vision
Fine-tune a pretrained classifier on a small dataset
Transfer-learning accuracy with calibration on a small public dataset (CIFAR-scale), CPU-sized.
Computer vision
Architecture claim at ImageNet scale
Train and compare architectures on ImageNet-1k under a matched recipe.
Needs a GPU you connect
Computer vision
Robustness under corruptions
Accuracy of pretrained models on CIFAR-10-C by corruption type and severity.
Computer vision
Train a detector or segmenter
Detection/segmentation model trained and evaluated on COCO-style data.
Needs a GPU you connect
Computer vision
Label-noise and test-set overfitting audit
Find mislabeled test items and compare accuracy on a re-collected test set (CIFAR-10.1).
Computer vision
Novel-view synthesis benchmark
NeRF/3D Gaussian splatting quality (PSNR/SSIM/LPIPS) on a standard scene set.
Needs a GPU you connect
Formal methods
Model-check a TLA+ specification
TLC run over a bounded instance: invariants hold or a counterexample trace.
Formal methods
Find models or counterexamples with Alloy
Alloy analysis of a relational model within a stated scope.
Formal methods
Symbolic execution of Python against its contracts
Counterexamples to a Python function's pre/postconditions.
Formal methods
Property-based testing with shrinking
Minimal failing input for a stated property of real code.
Formal methods
Machine-checked proof in Coq
A theorem stated and proved in Coq using the standard library.
Formal methods
Is this program transformation sound?
SMT encoding of a transformation with a proof or counterexample.
Human-computer interaction
Analyse my controlled user study
Mixed-effects analysis of a within-subjects experiment with effect sizes and CIs.
Human-computer interaction
Power analysis and preregistration
Required sample size and a preregistration document with the analysis script dry-run on simulated data.
Human-computer interaction
Fit Fitts' law to pointing data
Model fit, throughput and per-condition comparison from pointing trials.
Human-computer interaction
Code interview transcripts
Codebook, coded excerpts and inter-rater agreement for qualitative data the researcher provides.
Human-computer interaction
Analyse deployment logs
Usage patterns and retention from in-the-wild interaction logs.
Human-computer interaction
Survey instrument analysis
Reliability (alpha/omega) and factor structure of a questionnaire.
Human-computer interaction
Reproduce a paper from its OSF artifact
Re-run a published analysis from its OSF materials and compare numbers.
Information retrieval
BM25 ad-hoc evaluation on a test collection
nDCG/MAP/recall of BM25 on a public collection with qrels.
Information retrieval
Dense retrieval across the full BEIR suite
Zero-shot dense-retriever results on all BEIR datasets.
Needs a GPU you connect
Information retrieval
Statistically sound system comparison
Paired tests with multiple-comparison control over per-query metrics from run files.
Information retrieval
Retrieval-augmented generation evaluation
Retrieval and answer quality of a RAG pipeline on a QA benchmark.
Needs a GPU you connect
Information retrieval
Corpus and query-log analytics
SQL analytics over a large document or interaction corpus.
Machine learning
Is an improvement real or seed noise?
Seed-level significance test and effect size for a claimed gain.
Machine learning
Fit a scaling law
Scaling exponents with bootstrap intervals from a run table.
Machine learning
Ablation study of a small model
Causal ablation of components on a CPU-sized task with seeds and compute matched.
Machine learning
Budgeted hyperparameter search on tabular benchmarks
Tuned baselines on an OpenML suite with the search budget disclosed.
Machine learning
Compare architectures at GPU scale
Matched-recipe training runs of models too large for CPU.
Needs a GPU you connect
Machine learning
Language-model benchmark evaluation
lm-evaluation-harness scores for open models with the harness version pinned.
Needs a GPU you connect
Machine learning
Benchmark contamination audit
N-gram/near-duplicate overlap of a benchmark with a pretraining-corpus sample.
Machine learning
Reproduce a published result at CPU scale
Re-run a paper's code from GitHub and compare its numbers.
Natural language and LLMs
Tokenisation and perplexity experiments
Tokenizer behaviour and perplexity of small models on a corpus.
Natural language and LLMs
Score translations or summaries
BLEU, chrF and ROUGE with signatures for system outputs.
Natural language and LLMs
Fine-tune a small language model
Fine-tuned small model (≤ ~125M params) with evaluation against the base.
Natural language and LLMs
LoRA fine-tune of a billion-parameter model
Parameter-efficient fine-tuning of a 1B+ model with a disclosed search budget.
Needs a GPU you connect
Natural language and LLMs
Social-bias probing of open models
BBQ/CrowS-Pairs-style results disaggregated by group.
Needs a GPU you connect
Natural language and LLMs
Semantic search over a corpus
Embedding index and retrieval quality over a document set.
Natural language and LLMs
Text-classification baselines
TF-IDF/linear and small-transformer baselines on a public dataset with error analysis.
Natural language and LLMs
Serve a model for high-throughput inference
Throughput/latency of a served model under load (vLLM).
Needs a GPU you connect
Robotics and control
Simulate a robot and close a control loop
MuJoCo/PyBullet simulation with a feedback controller and tracking metrics.
Robotics and control
Classical control design
LQR, pole placement, Bode and root-locus analysis of a plant.
Robotics and control
State estimation with Kalman-family filters
EKF/UKF/particle-filter estimates and consistency checks on a dataset.
Robotics and control
Motion-planning benchmark
OMPL planners compared on success rate and path quality.
Robotics and control
Embodied navigation in photorealistic scenes
Habitat navigation agents on standard scene datasets.
Needs a GPU you connect
Robotics and control
Pose-graph SLAM back end
Least-squares pose-graph optimisation with loop closures on a dataset.
Security and cryptography
Find an input that reaches a code location
Symbolic execution of a binary to a target block (e.g. a crackme's success path).
Security and cryptography
Analyse a binary's structure
Sections, symbols, disassembly and emulation of an ELF binary.
Security and cryptography
Break a textbook cryptographic construction
Working attack (padding oracle, CBC bit-flip, small-exponent RSA) with explanation.
Security and cryptography
Vulnerability trends from NVD
CVE counts, CWE mix and CVSS trends for a product or class.
Security and cryptography
Scan a sample corpus with YARA
Rule matches and false-positive checks over files the researcher supplies.
Security and cryptography
Membership-inference evaluation
Attack AUC against a trained model and the effect of a defence.
Systems and databases
Analytical SQL on large Parquet data
Query results, plans and engine comparison over multi-GB columnar data.
Systems and databases
Query-optimiser study on TPC-H
Plan quality and regret across queries and scale factors.
Systems and databases
Deterministic simulation testing of a protocol
A distributed protocol run under a seeded scheduler until a correctness bug appears, with the replay.
Systems and databases
Find a tail-latency regression
Percentile analysis with CIs and change-point detection on latency measurements.
Systems and databases
Re-run a reproducibility package
A paper's artifact re-executed from Zenodo/GitHub with results compared.