# Propab > Propab is an AI co-researcher for ambitious science. It reviews the literature, analyses data, writes and runs code in ready scientific environments, runs computational experiments on your compute or GPUs, and keeps every result traced to its source, code and data. Propab is for working scientists in biology, chemistry, physics and computer science. A run can last minutes or days; it reads the literature (290M+ papers, patents, trials, grants), fetches data from 130+ databases, runs code in 40+ ready environments on a hosted sandbox or the researcher's own machines and GPUs, checks its results, and writes them up with every number traced to its source, code and data. Free to start; paid plans from $20 a month. ## Product - [Overview](https://propabai.com/index.md): what Propab does, how a run works, and the questions researchers ask first - [Pricing](https://propabai.com/pricing.md): Free, Pro, Max, Team and Enterprise, and what each includes - [About](https://propabai.com/about.md): what Propab is building and the rules it builds by ## Use cases - [Biology](https://propabai.com/use-cases/biology.md): real runs and the workflows ready for each subfield - [Chemistry](https://propabai.com/use-cases/chemistry.md): real runs and the workflows ready for each subfield - [Physics](https://propabai.com/use-cases/physics.md): real runs and the workflows ready for each subfield - [Computer science](https://propabai.com/use-cases/computer-science.md): real runs and the workflows ready for each subfield ## Policies - [Privacy policy](https://propabai.com/legal/privacy) - [Terms of service](https://propabai.com/legal/terms) ## Optional - [Everything above in one file](https://propabai.com/llms-full.txt) # Propab — an AI co-researcher for ambitious science Propab is an AI co-researcher for ambitious science. It reviews the literature, analyses data, writes and runs code in ready scientific environments, runs computational experiments on your compute or GPUs, and keeps every result traced to its source, code and data. ## Ready for the field you work in Environments, methods and data for biology, chemistry, physics and computer science are set up before you ask your first question: Experimental and applied, Bioinformatics, Single-cell, Population genetics, Genomics, Systems biology, Proteomics, Microbiology, Molecular biology, Structural biology, Biochemistry, Immunology, Cell biology, Microscopy, Neuroscience in biology, and the same depth in the other three fields. See the [use cases](https://propabai.com/use-cases.md). ## One space, from the first paper to the manuscript Research is usually spread across a dozen tools: search in one, download in another, convert with a script, run on a cluster, write it up somewhere else. In Propab it is one place: - **Literature** — 290M+ papers, plus patents, clinical trials and research grants. - **Data** — 130+ databases (UniProt, ChEMBL, Ensembl, PubChem, NCBI and more), Hugging Face and Kaggle datasets, GitHub. - **Files** — structures, sequences, single-cell data, images and tables, opened and converted as they are. - **Code** — 40+ ready environments and 350+ proven methods. - **Compute** — a hosted sandbox by default; your own machines over SSH and your GPU accounts when the work needs more. - **Writing** — citations kept, a Zotero library connected, manuscripts exported to Word, LaTeX and Markdown. ## What it does - **Every result, traced and reproducible** — each claim opens to its source, the code that ran and the data it used. - **Real work, at scale** — thousands of experiments over days, not a paragraph in a minute. - **Runs on your compute** — your laptop, cluster or GPUs; Propab sets up, scales up and gives back what it took. - **Works with your data** — your files and formats, and the databases your lab already relies on. - **You stay in charge** — steer, pause or redirect a run at any point; Propab asks when the call is yours. ## How a run works 1. Your question, in your words. 2. Propab reads the field. 3. It plans the work. 4. It runs experiments. 5. It checks the result. 6. It writes it up, and goes round again until the question is answered. ## Questions researchers ask ### How is Propab different from asking a chat assistant? A chat assistant answers from what it remembers. Propab goes and does the work: it searches the literature, downloads the data, writes and runs the code, checks the result, and keeps going until the question is answered. A run can last minutes or days, and everything it made is kept. ### How do I check that a result is right? Every number, figure and claim in a run opens to how it was made: the source passage it came from, the code that produced it, the inputs it read and the output it printed. Propab also runs its own checks, and it tells you when something is uncertain rather than smoothing it over. ### Does it work in my field? Propab has ready environments for more than forty subfields across biology, chemistry, physics and computer science, each with its tools installed and tested, plus hundreds of methods it knows how to use. If your work sits between fields, it combines them. ### Can I use my own data? Yes. Upload files or point Propab at the databases and repositories you already use. It reads common scientific formats as they are, such as structures, sequences, single-cell data, images and tables, and converts between them when a step needs it. ### Where does the code run? In a sandbox Propab manages for you by default. When a problem needs more, connect your own machine or cluster over SSH, or your GPU account, and Propab will set up the environment there, run the work, and release what it took when it is done. ### Can I stop or change direction mid-run? Yes. You can send a message at any point to add a constraint, change the plan or ask a question, and Propab adapts without starting over. You can pause a run and continue it later, and it asks you when a decision is genuinely yours to make. ### Who can see my runs? Your runs are private to you. Nothing is public unless you share it: on paid plans any run can become a read-only replay with private details removed, and you can turn the link off at any time. ### Can Propab help me write the paper? It keeps every source it read, can pull in your Zotero library, and writes up results with citations. Manuscripts export to Word, LaTeX (ready for Overleaf) and Markdown, and every number in them still points to the data it came from. ### Which models does it use? Leading frontier models. On Auto, Propab picks the right one for each step; on paid plans you can choose any model yourself. The work around the model, including the environments, data, compute and checks, is what makes the results hold up. Links: [Pricing](https://propabai.com/pricing.md) · [About](https://propabai.com/about.md) · [Privacy](https://propabai.com/legal/privacy) · [Terms](https://propabai.com/legal/terms) --- # Propab pricing Every plan reads the full literature, runs in ready environments and keeps every result traced. Paid plans add model choice, more usage and sharing. Paying yearly costs eleven months instead of twelve. ### Free — $0, always To try Propab on a real question. - Auto mode picks the model for each step - Full literature search - Every ready environment, database and method - Download every artifact and export manuscripts - Connect your own machines and GPUs - A hosted workspace: runs keep going with your laptop closed* - You own the IP of everything you generate ### Pro · Standard — $20/month or $220/year Good for day-to-day research. - Everything in Free, plus: - Choose any model, or leave it on Auto - More usage, for everyday work - Share runs as public replays - Several tasks running at once - More memory and disk for each run - Email support ### Pro · Plus — $60/month or $660/year For researchers who work with Propab all day. - Everything in Free, plus: - Choose any model, or leave it on Auto - More usage, for everyday work - Share runs as public replays - Several tasks running at once - More memory and disk for each run - Email support ### Max · Power — $100/month or $1,100/year For power users and days-long runs. - Everything in Pro, plus: - Our highest usage - The most tasks running at once - The most memory and disk for each run - Priority for large compute jobs - Early access to new models and features - Priority support ### Max · Highest — $200/month or $2,200/year The most Propab can do, for the heaviest work. - Everything in Pro, plus: - Our highest usage - The most tasks running at once - The most memory and disk for each run - Priority for large compute jobs - Early access to new models and features - Priority support ### Team Pro — $20 per seat per month · Team Max — $100 per seat per month Pro or Max for every researcher, with shared admin, usage by member and one invoice. ### Enterprise — custom pricing Flexible deployment (our cloud, a dedicated instance, or your own cloud account); single sign-on, admin controls and audit logs; your internal data, instruments and tools connected; compute sized to your programme; direct support. \* Runs on machines you connect yourself need those machines to stay on. ## Questions about plans ### What happens when I reach my limit? The run pauses where it is, with everything it has done kept. It picks up again on its own when your usage resets, or straight away if you move to a larger plan. ### What counts as usage? The model work Propab does for you: reading, planning, writing code and checking results. Searching the literature does not count, on any plan. ### Which model does Propab use? On Free, Propab uses Auto and picks the model for each step. On paid plans you can choose any model for a run, or leave it on Auto. ### Who owns what Propab makes? You do. You own the IP of everything you generate: data, code, figures, results and manuscripts, on every plan. ### Do I need to keep my computer on? No. Your workspace is hosted, so a run keeps going with your laptop closed and you can check in from anywhere. The one exception is a machine you connect yourself, such as your own server or cluster: it has to stay on while Propab uses it. ### Can I share a run? On paid plans, any run can be shared as a read-only replay with private details removed. You can turn the link off at any time. ### Do I pay for GPUs through Propab? No. Connect your own GPU account or machine and your provider bills you directly. Propab starts the machine, runs the work and shuts it down when it is done. ### Is there a discount for a year? Yes. Paying yearly costs eleven months instead of twelve. --- # About Propab Propab is an AI co-researcher for ambitious science. It reviews the literature, analyses data, writes and runs code in ready scientific environments, runs computational experiments on your compute or GPUs, and keeps every result traced to its source, code and data. Propab is building the research infrastructure that lets an AI co-researcher work on a real problem for days, with the tools, data and compute a lab would use, and with every result traced. It was founded by Shani Singh. ## What we believe 1. **Intelligence is not the bottleneck. The work around it is.** Models can reason about science; what stops them is a sandbox that dies, a dataset that will not download, a tool nobody installed, a check nobody ran. 2. **A run ends when the question is answered.** Propab can work on one problem for days and run thousands of experiments, never cut short by a step counter. Limits apply to money, compute and time. 3. **Every result must be checkable.** Every number opens to the code that produced it, the data it read and the sources it relied on. ## How we build Correctness first · Nothing made up · Built for the real size of the work · The interface is the science. --- # Propab use cases Real runs, the questions exactly as they were asked, and the workflows ready for each subfield. - [Biology](https://propabai.com/use-cases/biology.md) — Experimental and applied, Bioinformatics, Single-cell, Population genetics, Genomics, Systems biology, Proteomics, Microbiology, Molecular biology, Structural biology, Biochemistry, Immunology, Cell biology, Microscopy, Neuroscience - [Chemistry](https://propabai.com/use-cases/chemistry.md) — Analytical chemistry, Cheminformatics, Drug discovery, Materials chemistry, Molecular dynamics, Process engineering, Quantum chemistry, Reactions and synthesis - [Physics](https://propabai.com/use-cases/physics.md) — Astrophysics, Computational physics, Condensed matter, Fluid and continuum mechanics, Nuclear physics, Optics and electromagnetism, Particle physics, Plasma physics, Quantum physics, Relativity and mathematical physics, Statistical mechanics - [Computer science](https://propabai.com/use-cases/computer-science.md) — Algorithms and theory, Compilers and languages, Computer architecture, Computer vision, Formal methods, Human-computer interaction, Information retrieval, Machine learning, Natural language and LLMs, Robotics and control, Security and cryptography, Systems and databases --- # Biology with Propab Differential expression, single-cell, structures, variants, trials and the literature behind them, worked through end to end, with every step you can open. ## Real runs - **Genomics** — [replay](https://propabai.com/replay/share_h1s8bFWvG58GslT9dUy_gERxjXYmfjFk): “Take a public bulk RNA-seq study from GEO comparing treated and control samples in a human cell line (pick one with raw counts available), run a proper differential expression analysis with DESeq2 or an equivalent, and follow it with pathway enrichment. Deliver a volcano plot, a heatmap of the top genes, the ranked gene table and a short report that says which pathways move and how confident we should be.” Made: Volcano plot, Heatmap, Report. - **Structural biology** — [replay](https://propabai.com/replay/share_LztRLxSuBAPeXcQCNXoR_uxF4SsOB36l): “End-to-end computational study of KRAS G12C. Do real computation, not mock data. Every number you report must come from code you ran. If a step fails or a tool can’t be installed, say so and use the closest working alternative. Fetch human KRAS, HRAS, NRAS and 10–15 RAS orthologs across species, align them, score per-residue conservation and build the phylogenetic tree…” Made: Alignment, Phylogeny, 3D structures. ## Workflows ready today ### Experimental and applied - **Preclinical literature review** — Searches the literature on a target or drug and extracts in-vitro and in-vivo findings into a table, one source per row. - **Clinical-trial landscape** — Maps registered trials for a disease by mechanism, phase, status and sponsor, with counts and a timeline. - **Target–disease evidence** — Pulls Open Targets association scores and the evidence behind them for a target or disease, by evidence type. - **Disease progression from longitudinal omics** — Models how molecular profiles change over time in patients and groups patients by trajectory. - **Experimental design and power** — Computes power and sample size, lays out batches and fixes the multiple-testing plan before data are collected. - **ADMET prediction and docking** — Predicts ADMET liabilities and docks compounds into a target pocket with a re-docking control. - **Biomarker panel selection** — Selects a minimal biomarker panel with LASSO inside cross-validation and reports held-out performance. - **Survival analysis** — Kaplan–Meier curves, Cox regression and risk groups for a cohort. - **Signature reversal (connectivity map)** — Finds compounds whose LINCS L1000 signatures oppose a disease signature. - **Literature review** — Searches and synthesises the literature on a target, disease or question, with a source for every claim. - **Trials for a target** — Lists ClinicalTrials.gov records for drugs against a target, with phase, status, sponsor and NCT id. - **Report and figure deliverables** — Builds the run's report as PDF, Word or PowerPoint with publication-style figures. - **Tractability and druggability** — Reports Open Targets tractability buckets and detects pockets on the target structure. - **Drug repurposing** — Proposes new indications for existing drugs from shared targets, network proximity and signature reversal. - **Adverse-event signals** — Computes disproportionality (PRR/ROR) signals for a drug from FDA adverse-event reports. - **Virtual screening** — Docks a compound library against a target, triages the hits and summarises structure–activity. - **Off-target and safety pharmacology** — Predicts secondary-pharmacology, hERG and selectivity liabilities for compounds. - **Real-world evidence** — Outcomes analysis on an EHR or claims extract the researcher provides. - **Clinical trial design and simulation** — Simulates adaptive or enrichment designs to compare power and error rates. - **Meta-analysis** — Pools effect sizes across studies with random-effects models and reports heterogeneity. - **Microplate layout** — Plate layouts with randomisation, edge-effect handling and covariate balance. ### Bioinformatics - **Find public omics datasets** — Finds and catalogues public datasets that fit a question across GEO, ENA/SRA, ArrayExpress, PRIDE, MetaboLights, CELLxGENE and GDC, and says which are usable. - **Pathway enrichment (GSEA / ORA)** — Runs GSEA or over-representation on a gene list against MSigDB, GO and Reactome sets. - **Methods landscape review** — Compares tools and pipelines for a task from published benchmarks. - **Phylogenetic tree from sequences** — Aligns homologous sequences, trims the alignment and infers a maximum-likelihood tree with support values. - **Homology search and domain annotation** — Finds homologues with BLAST and HMMER against Swiss-Prot and Pfam and annotates protein domains. ### Single-cell - **Target prioritisation from single-cell data** — Finds the cell types that change in a disease in public single-cell data and ranks genes by that signal plus human genetic evidence. - **scRNA-seq clustering and cell-type annotation** — Takes a count matrix through QC, normalisation, batch integration, clustering and marker-based cell-type labels. - **Trajectories, pseudotime and RNA velocity** — Orders cells along a differentiation process, finds genes that change along it, and estimates RNA velocity where spliced/unspliced counts exist. - **Cell–cell communication** — Infers ligand–receptor interactions between annotated cell types. - **Transcription-factor activity and regulatory networks** — Scores each TF's regulon activity per cell and infers TF–target co-expression networks. - **Perturb-seq / CROP-seq screen analysis** — Assigns guides to cells, removes unperturbed cells and measures each knockdown's transcriptional effect. - **Spatial transcriptomics (Visium)** — QC, spatial domains and neighbourhood enrichment between cell types on Visium data. - **Virtual perturbation prediction** — Predicts single-cell responses to unseen genetic perturbations with a foundation model and scores them on held-out data. (needs a GPU you connect) - **scATAC, multiome and CITE-seq** — Chromatin-accessibility and multimodal single-cell analysis: peaks, clustering, gene activity and protein–RNA integration. ### Population genetics - **GWAS to genes (TWAS)** — Tests which genes' predicted expression is associated with a trait, per tissue, from GWAS summary statistics. - **Fine-mapping GWAS loci** — Gives each variant at a GWAS locus a probability of being causal, with credible sets. - **Mendelian randomisation** — Estimates whether an exposure causally affects an outcome from two GWAS, with IVW, MR-Egger, weighted median and sensitivity checks. - **Polygenic risk scores** — Applies a published PGS Catalog score to genotypes and compares score distributions across populations. - **Population structure and selection scans** — Genotype PCA, population structure and selection statistics from population genotypes. ### Genomics - **Variant annotation** — Annotates a VCF with consequence, population frequency and clinical significance, and flags likely damaging variants. - **Bulk RNA-seq differential expression** — Tests genes for differential expression between conditions from a count matrix and returns fold changes, adjusted p-values and plots. - **Upstream-regulator analysis** — Finds the transcription factors most likely driving an observed differential-expression signature. - **Genetic constraint gating** — Flags loss-of-function-intolerant genes (pLI, LOEUF) in a gene list. - **Cancer cohort genomics** — Mutation and copy-number frequencies across TCGA and MSK cohorts, with mutant-versus-wild-type comparisons. - **Consensus disease signature** — Combines several public expression studies into one up/down disease signature by meta-analysis. - **Tissue expression and safety** — Shows where a target is expressed across normal tissues and cell types. - **RNA-seq quantification from FASTQ** — Pseudo-aligns raw RNA-seq reads to a transcriptome and produces gene counts. ### Systems biology - **Gene co-expression modules** — Finds co-expressed gene modules and their hub genes and relates modules to sample traits. - **Multi-omics factor analysis (MOFA+)** — Finds shared and layer-specific latent factors across two or more omics layers measured on the same samples. - **Multi-omics sample clustering** — Clusters samples across transcriptome, proteome and metabolome and checks cluster stability. - **Knowledge-graph target ranking** — Ranks targets by multi-hop evidence paths through a biomedical knowledge graph. - **Gene essentiality on a metabolic model** — Knocks out every gene in a genome-scale metabolic model on chosen media and reports which genes are essential and which only on one medium. - **Kinetic model simulation** — Loads a curated kinetic model from BioModels and runs time courses, steady states and parameter scans. ### Proteomics - **Proteomics differential abundance** — Tests proteins for differential abundance from a quantified mass-spec table, with FDR control. ### Microbiology - **Microbiome community analysis** — Compares alpha and beta diversity and differential abundance across groups from a taxon count table. - **Bacterial genome assembly and gene calling** — Trims reads, assembles an isolate genome, predicts genes and reports assembly statistics. - **Shotgun read classification against a chosen panel** — Builds a Kraken2 index from the genomes you choose and classifies metagenomic reads against it. ### Molecular biology - **PCR / qPCR primer design** — Designs primer pairs and TaqMan probes for a target and checks dimers, hairpins and off-target hits. - **Restriction mapping** — Digests a sequence in silico and predicts fragment sizes and the gel. ### Structural biology - **Protein structure prediction** — Predicts a protein or complex structure with per-residue confidence. (needs a GPU you connect) - **De-novo binder and antibody design** — Generates binder backbones against a target site, designs sequences and filters by predicted fold and interface. (needs a GPU you connect) - **Ligand binding-mode analysis** — Picks the best co-crystal structure and lists the pocket residues that contact the ligand. - **Targeted protein degradation design** — Models ternary complexes and proposes PROTAC linkers. (needs a GPU you connect) - **Predicted versus experimental structure** — Superposes an AlphaFold model onto the experimental structure and reports error by region against the model's confidence. - **Protein molecular dynamics** — Solvates, equilibrates and simulates a protein and analyses its stability and flexibility. (needs a GPU you connect) ### Biochemistry - **Binding-affinity model** — Trains and validates a model that predicts compound–target affinity from public measurements, with scaffold and cold-target splits. - **Drug potency and selectivity profile** — Pulls IC50, Ki and Kd for a drug against its target and its off-targets. - **Enzyme kinetics fitting** — Fits Michaelis–Menten and inhibition models to rate data and compares them. ### Immunology - **Cell-therapy product QC** — Single-cell QC scorecard for a cell-therapy product: viability, composition, activation and exhaustion states. - **Immune repertoire (TCR/BCR)** — Clonotypes, clonality and diversity from AIRR-formatted repertoires. - **Immune-cell deconvolution** — Estimates immune-cell fractions in bulk expression samples. ### Cell biology - **Flow cytometry analysis** — Gates or clusters cytometry events and reports population frequencies. ### Microscopy - **Digital pathology** — Tiles whole-slide images and trains slide-level classifiers. (needs a GPU you connect) - **Nuclei segmentation and measurement** — Segments nuclei or cells in fluorescence images and measures size, shape and intensity per object. ### Neuroscience - **Intrinsic electrophysiology from patch-clamp** — Extracts rheobase, f–I slope, resting potential, input resistance and membrane time constant from current-clamp sweeps. --- # Chemistry with Propab Compound data, docking, quantum chemistry, molecular dynamics and materials screens, run at the scale the question needs and checked against what is known. ## Real runs - **Cheminformatics** — [replay](https://propabai.com/replay/share_a_cIvoLdi-04i_7nZVYPu0PdFt1YfMYt): “For EGFR, pull measured inhibitors from ChEMBL, filter them by drug-likeness, cluster them by scaffold, and pick 20 diverse candidates for testing with a rationale. Deliver a chemical-space plot, the candidate table and a short report.” Made: Chemical space, Candidate table, Report. ## Workflows ready today ### Analytical chemistry - **Identify an unknown from its EI mass spectrum** — Library search of a 70 eV EI spectrum with a calibrated accept/refuse threshold, so 'not in the library' is a possible answer. - **Process an NMR FID into a quantitative spectrum** — Phased, referenced 1D spectrum with an integral table; qNMR purity against an internal standard. - **Quantify an analyte by chromatography** — Integrated peaks, internal-standard calibration, concentration and expanded uncertainty from raw traces. - **Resolve two co-eluting peaks** — Deconvolved areas of overlapping peaks and whether the separation is adequate. - **Molecular formula from accurate mass** — Candidate formulas ranked by mass error and isotope-pattern fit. - **Untargeted LC-MS feature table** — Detected, aligned and annotated features across an LC-MS run set, with spectral-library matches. - **Multivariate calibration of spectra** — PLS calibration or PCA of a spectral set, with cross-validated error and outliers flagged. - **Validate an analytical method** — Linearity, accuracy, precision, LOD/LOQ and a GUM uncertainty budget to ICH Q2 / Eurachem. ### Cheminformatics - **Build a modelling table from ChEMBL** — All measured activities for a target, curated to one value per compound with units and assay types reconciled. - **Standardise and deduplicate a library** — Canonical, salt-stripped, deduplicated structures with a log of every change. - **Similarity screen against known actives** — Library ranked by fingerprint similarity to actives, with an honest check of whether the ranking beats trivial property baselines. - **QSAR model with a scaffold split** — Property/activity model validated on held-out scaffolds, with an applicability domain and a scrambled-label control. - **SAR table and matched molecular pairs** — R-group decomposition of a series and the potency change each transformation buys. - **Map chemical space and pick a diverse subset** — 2D map (UMAP/t-SNE) of a library and a maximally diverse purchase list. ### Drug discovery - **Dock a library and triage the hits** — Every compound docked into a site, scored, and a shortlist with liabilities named and the score's limits stated. - **Validate a docking protocol by re-docking** — RMSD of the re-docked co-crystal ligand to its crystal pose before any score is trusted. - **Prepare a receptor from a crystal structure** — Protonated, repaired receptor with waters/cofactors decided explicitly, ready for docking. - **Find and rank binding pockets** — Cavities on a protein with volume and druggability descriptors. - **Protein-ligand interaction fingerprints** — Which H-bonds, salt bridges and contacts each pose makes, compared across a series. - **Collect every measured compound for a target** — SAR table of all ChEMBL/BindingDB measurements for a target, deduplicated and unit-normalised. - **Relative binding free energies for a series** — FEP ranking of congeneric ligands with cycle-closure errors. (needs a GPU you connect) - **Co-fold a target with a ligand** — Complex structure and predicted affinity from Boltz-2 or Chai-1 without a crystal structure. (needs a GPU you connect) - **Selectivity and off-target profile** — Activity of a series across a target family, and likely off-targets by ligand similarity. - **Lead-optimisation loop** — Rounds of analogue enumeration, filtering, docking and scoring with the reasoning per round. ### Materials chemistry - **Is this compound stable? (convex hull)** — Formation energy and energy above the hull against all competing phases, with the phase set's completeness checked. - **Space group and structure matching** — Space group, Wyckoff positions, standardised cell, and whether two structures are the same material. - **Elastic constants and moduli** — Elastic tensor, bulk/shear/Young's moduli and mechanical stability criteria. - **Identify a phase from its XRD pattern** — Simulated powder patterns of candidate phases matched against a measured pattern. - **ML property model for materials** — Featurised structures (matminer) and a validated model for a target property. ### Molecular dynamics - **Build and equilibrate a solvated protein** — Solvated, neutralised periodic system from a PDB entry, equilibrated NVT/NPT with convergence checks. - **Transport properties from equilibrium MD** — Self-diffusion coefficient (with finite-size correction) and viscosity of a liquid. - **Analyse a trajectory** — RMSD, RMSF, radius of gyration, H-bonds, contacts and RDFs of a supplied trajectory. - **Production protein MD at microsecond scale** — Long unbiased trajectories of a protein for conformational analysis. (needs a GPU you connect) - **Free-energy surface by enhanced sampling** — Metadynamics or umbrella sampling along a collective variable, with MBAR/WHAM and convergence evidence. - **Markov state model from many trajectories** — Metastable states, implied timescales and transition rates. ### Process engineering - **Design a distillation separation** — Property model choice, VLE curve at column pressure and a sized column (Rmin, stages, feed stage), including refusing a split an azeotrope forbids. - **Thermodynamic cycle with real-fluid properties** — State points, efficiency and COP of a power or refrigeration cycle. - **Size a heat exchanger** — Area, pressure drop and rating of a shell-and-tube exchanger for a duty. - **Reactor with detailed kinetics** — Species/temperature profiles and sensitivity analysis from a published CHEMKIN mechanism. - **Flame speed and ignition delay against experiment** — Laminar burning velocity and ignition delay vs equivalence ratio, compared with measured data. - **Regress binary interaction parameters** — Activity-model parameters fitted to VLE measurements, with the fit's residuals. - **Converge a flowsheet with recycles** — Rigorous column and recycle convergence into a stream table. - **Techno-economic analysis** — Capital and operating cost, NPV and minimum selling price with a sensitivity tornado. - **Process dynamics and controller tuning** — Dynamic response of a unit and PID tuning with step tests. ### Quantum chemistry - **Electronic energy at a stated level of theory** — HF/DFT/MP2 energy of a molecule in named basis sets. - **Optimise a geometry and confirm the stationary point** — Optimised structure, bond lengths/angles, and a Hessian proving a minimum (0 imaginary) or saddle (exactly 1). - **Transition state, IRC and barrier** — Located TS, IRC connecting it to both minima, and the barrier. - **Conformer ensemble with Boltzmann weights** — Conformers generated, optimised (GFN2 then DFT) and Boltzmann-weighted. - **Non-covalent interaction energy** — Counterpoise-corrected interaction energies benchmarked against S22/S66 references. - **UV-Vis spectrum by TDDFT** — Vertical excitations and oscillator strengths, broadened into a spectrum. - **High-accuracy or large-system calculation** — CCSD(T) or large-basis DFT on systems beyond a few dozen atoms. (needs a GPU you connect) ### Reactions and synthesis - **One retrosynthetic step** — Ranked precursor sets for a target from a template library, with a null-model check. - **Triage a library by synthesisability** — SA/SC scores and a ranked, filtered list. - **Yield model from public HTE data** — Yield-prediction model trained on an open high-throughput dataset (e.g. Buchwald-Hartwig), with uncertainty. - **Design the next batch of condition experiments** — Bayesian-optimisation suggestion of the next conditions to run from results so far. - **Enumerate a combinatorial library** — Products of building blocks under reaction templates, as a reopenable SDF. - **Compare routes on green metrics** — PMI, E-factor and atom economy of candidate routes. --- # Physics with Propab Simulations, observational data and theory worked through carefully, with error bars, finite-size checks and the literature value beside every result. ## Real runs - **Statistical mechanics** — [replay](https://propabai.com/replay/share_UCErCeW5OVnexzEVKJNSWL3WbZz6OTmk): “Reproduce the 2D Ising model ferromagnetic phase transition from first principles. Run Monte Carlo (Metropolis, plus a Wolff cluster algorithm near criticality) on square lattices of several sizes, measure magnetisation, susceptibility, specific heat and the Binder cumulant, and estimate the critical temperature by finite-size scaling. Compare your estimate with Onsager's exact result and with values reported in the literature.” Made: Observables, Finite-size fit, PDF report. ## Workflows ready today ### Astrophysics - **Known-exoplanet population study** — Mass-radius and period-radius diagrams (e.g. the radius valley) from the NASA Exoplanet Archive, with the selection cuts stated. - **Redshifts and emission lines from SDSS spectra** — Redshift, line fluxes and a BPT classification for a galaxy sample, from SDSS line catalogues or the spectra themselves. - **Hubble constant from Pantheon+ supernovae** — H0 / Omega_m posterior from the Pantheon+SH0ES distance moduli with the full covariance matrix, with a corner plot. - **Gravitational-wave population from published posteriors** — Mass and spin distribution of merging black holes by hierarchical inference over the GWTC posterior samples. - **Analyse my cosmological simulation output** — Halo profiles, projections and power spectra from a simulation snapshot the researcher uploads. ### Computational physics - **Check an integrator's order and conservation** — Measured convergence order and energy/momentum drift of an integrator (symplectic vs Runge-Kutta) on a test problem. - **Long-term stability of a planetary system** — MEGNO / Lyapunov map and survival time of a multi-planet system over millions of orbits. - **Solve a PDE and prove the convergence order** — Finite-difference solution of a heat/wave/Poisson problem with a grid-refinement study against the exact solution. - **Monte Carlo estimate with an honest error bar** — Monte Carlo integral or expectation with variance reduction and a convergence plot. - **Spectrum of my time series** — Power spectral density (Welch/multitaper), peaks and their significance for a signal the researcher supplies. - **Derive equations of motion and simulate them** — Lagrangian to equations of motion symbolically, then integrated and plotted (phase portraits, energy). - **Fit a simulation's parameters by gradient** — Differentiable simulation fitted to data by automatic differentiation. - **Solve a stiff ODE or DAE system** — Stiff kinetics or an index-reduced DAE solved with an implicit integrator, with a solver-tolerance study. ### Condensed matter - **Tight-binding bands and density of states** — Band structure and DOS of a lattice model, checked against the analytic dispersion. - **Topological invariant of a model** — Chern number or Z2 invariant of a Hamiltonian, with the Berry-curvature or Wilson-loop plot. - **Ground state of an interacting 1D model by DMRG** — Energy, correlations and entanglement entropy of a spin or Hubbard chain with TeNPy or quimb. - **Exact diagonalisation of a small interacting cluster** — Spectrum and correlation functions of a Heisenberg or Hubbard cluster by Lanczos. - **Screen a materials database for a property** — Candidates ranked by band gap, stability or magnetism from open DFT databases. - **Classical spin Monte Carlo for a magnetic transition** — Ordering temperature and susceptibility of a Heisenberg/Ising magnet on a lattice. ### Fluid and continuum mechanics - **2D incompressible flow benchmark** — Lid-driven cavity or flow past a cylinder with drag, lift and Strouhal number against published reference values. - **Linear stability of a shear flow** — Orr-Sommerfeld eigenvalues, neutral curve and critical Reynolds number. - **Structural finite-element check against a reference** — Stress and deflection of a beam/plate/solid against an analytic or NAFEMS benchmark value. - **Topology optimisation of a structure** — Minimum-compliance layout at a volume fraction (SIMP), with a mesh-independence check. - **Bifurcations and chaos in a low-dimensional system** — Bifurcation diagram, Lyapunov exponent and period-doubling constants. - **Recover governing equations from data** — Sparse identification (SINDy) of a dynamical system from measured trajectories, with a noise-robustness test. ### Nuclear physics - **Decay an inventory forward in time** — Activities, decay heat and gamma lines of a nuclide inventory over time. - **Binding energies against a mass model** — AME2020 binding energies vs the semi-empirical mass formula, with shell-closure residuals. - **Measured cross sections against evaluated libraries** — EXFOR measurements plotted and compared against ENDF/B-VIII and JEFF evaluations for a reaction. - **Nucleosynthesis reaction network** — Energy generation and abundance flow for a burning stage from a REACLIB network. - **Astrophysical reaction rate from a cross section** — Thermonuclear rate vs temperature from an S-factor or narrow resonances, with its uncertainty. - **Analyse my gamma-ray spectrum** — Peak fits, energy/efficiency calibration and nuclide identification for a detector spectrum. ### Optics and electromagnetism - **Mie scattering of particles** — Extinction, scattering and backscatter efficiencies vs size parameter and wavelength. - **Thin-film stack reflectance** — Reflectance/transmittance spectra of a multilayer coating from real dispersion data. - **Diffraction pattern or telescope PSF** — Fraunhofer/Fresnel propagation of an aperture and the PSF with aberrations. - **Resonator modes and Q factor (2D FDTD)** — Cavity resonances, Q factors and transmission of a 2D photonic structure. - **Photonic-crystal band structure** — Band diagram and band gaps of a photonic crystal slab. - **Grating diffraction efficiency (RCWA)** — Diffraction efficiencies of a periodic grating vs wavelength and angle. - **Lens design and ray tracing** — Sequential ray trace of a lens with real glasses: spot diagrams, aberrations and a tolerance check. - **Analyse my measured S-parameters** — Touchstone network data de-embedded, fitted and plotted (Smith chart, return loss). ### Particle physics - **Rediscover the Z and J/psi in collision data** — Dimuon invariant-mass spectrum from CMS open data with fitted resonance masses and widths. - **Bump hunt with local and global significance** — Search for a resonance in a mass spectrum with a look-elsewhere-corrected significance. - **Upper limits with a HistFactory model** — 95% CL CLs limits on a signal from binned templates. - **Jet clustering and substructure** — Anti-kt jets and substructure observables of boosted objects on open data or supplied events. - **Standard Model prediction vs the world average** — Literature synthesis of a precision observable (e.g. muon g-2): each value, source and tension in sigma. - **Efficiency and scale factor by tag-and-probe** — Reconstruction/ID efficiency vs pT/eta in data and simulation, and their ratio. ### Plasma physics - **Characterise a plasma regime** — Debye length, plasma frequency, beta, collisionality and the regime they imply, for stated conditions. - **Kinetic dispersion relation** — Roots of the plasma dispersion function (Landau damping, ion-acoustic waves) against analytic limits. - **1D electrostatic particle-in-cell** — Two-stream instability growth rate from a PIC run against linear theory. - **Solar-wind driving of geomagnetic storms** — OMNI solar-wind coupling functions against Dst/AE over a storm set, with lag correlations. - **Turbulence spectrum from spacecraft magnetometer data** — Spectral index and break frequency of solar-wind or magnetosheath turbulence (Wind, MMS). - **Free-boundary tokamak equilibrium** — Grad-Shafranov equilibrium and the coil currents that produce a target shape. - **MHD benchmark problem** — Brio-Wu shock tube or Orszag-Tang vortex with a finite-volume MHD solver, checked against reference solutions and div B. - **Analyse my Langmuir-probe sweep** — Electron temperature, density and floating/plasma potential from an I-V characteristic. ### Quantum physics - **Decoherence of a driven qubit** — T1/T2 and Rabi dynamics from a Lindblad master equation. - **Variational algorithm on a spin Hamiltonian** — VQE or QAOA with gradients, converged energy and comparison with exact diagonalisation. - **Noisy circuit simulation at scale** — Circuits under a device noise model, up to the memory limit of a statevector/density matrix. - **Surface-code error-correction threshold** — Logical error rate vs physical error rate across code distances, and the crossing point. - **Tomography and randomised benchmarking** — State/process tomography and RB decay on a simulated device. - **Bound a Bell scenario by semidefinite programming** — NPA-hierarchy bound on a Bell expression with the SDP certificate. - **Tensor-network simulation of many qubits** — MPS/tensor-network simulation of a low-entanglement circuit beyond statevector size. - **Molecular Hamiltonian and VQE** — Qubit Hamiltonian of a small molecule and a VQE energy against full CI. - **Optimal control of a quantum gate** — GRAPE/CRAB pulse that implements a gate at target fidelity under constraints. ### Relativity and mathematical physics - **Curvature of a metric** — Christoffel, Riemann, Ricci and Einstein tensors, and whether the metric solves the field equations. - **Geodesics in a curved spacetime** — Orbits, perihelion precession, light bending or a black-hole shadow by geodesic integration. - **Random-matrix spectral statistics** — Level-spacing distribution and spectral form factor against GOE/GUE predictions. - **Asymptotic expansion of an integral or special function** — Series/asymptotic forms derived symbolically and checked numerically to high precision. - **Small lattice gauge simulation** — Pure-gauge U(1)/SU(2) Monte Carlo with Wilson loops and the string tension. - **Bethe ansatz vs exact diagonalisation** — Bethe-ansatz energies of an integrable chain checked against exact diagonalisation. ### Statistical mechanics - **Critical point and exponent by finite-size scaling** — Critical temperature and an order-parameter exponent with statistical and systematic errors separated. - **Cluster Monte Carlo that beats critical slowing down** — Wolff/Swendsen-Wang updates with measured autocorrelation times vs Metropolis. - **Density of states by Wang-Landau** — g(E) of a lattice model and thermodynamics at any temperature from one run. - **Error bars for correlated Monte Carlo data** — Integrated autocorrelation time and binning/blocking error for a Monte Carlo series. - **Percolation threshold and cluster scaling** — p_c and cluster-size exponents by finite-size scaling. - **Lennard-Jones fluid against NIST reference data** — Energy and pressure of LJ configurations compared with the NIST Standard Reference Simulation values. - **Stochastic-process simulation** — Exclusion processes, kinetic Monte Carlo or Gillespie simulations with scaling exponents. --- # Computer science with Propab Training runs, ablations, benchmarks and algorithm analysis, set up properly, run to completion and reported with the curves and tables behind them. ## Real runs - **Machine learning** — [replay](https://propabai.com/replay/share_A1gL8yr1_0wcY9GpqhQGr4ZaRoe4BgZk): “On a small transformer trained on a character-level language modelling task, ablate positional encodings, layer norm placement and dropout one at a time, and measure the effect on validation loss. Deliver the ablation table, curves and a short report.” Made: Ablation table, Loss curves, Report. ## Workflows ready today ### Algorithms and theory - **Empirical scaling of a solver on an instance family** — Where a problem family becomes intractable on this machine, with fitted scaling across solvers and the theory that explains it. - **Graph-algorithm study on generated families** — Runtime and solution quality of graph algorithms across generated graph families and sizes. - **SMT solver benchmarking with validated answers** — Solver comparison on SMT-LIB with every model and unsat core checked. - **Symbolic computation behind a proof** — Closed forms, recurrences and inequalities checked symbolically to support an argument. - **Anytime evaluation of a heuristic optimiser** — Quality-vs-time curves of heuristics on TSPLIB-style instances against known optima. ### Compilers and languages - **Is this peephole rewrite sound?** — An SMT proof or counterexample that an optimisation preserves semantics for all inputs. - **Parser, grammar and AST work** — Grammar written and tested, or ASTs of real code analysed. - **LLVM IR and optimisation-pipeline study** — IR built and run through optimisation passes, with before/after IR and JIT results. ### Computer architecture - **Roofline model of a kernel** — Arithmetic intensity vs attainable performance for a kernel on a named CPU/GPU from its datasheet numbers. ### Computer vision - **Score detections against ground truth** — COCO-style mAP/AR for a set of detections the researcher has. - **Classical image processing and matching** — Feature detection, matching, homography and geometric correction on images. - **Fine-tune a pretrained classifier on a small dataset** — Transfer-learning accuracy with calibration on a small public dataset (CIFAR-scale), CPU-sized. - **Architecture claim at ImageNet scale** — Train and compare architectures on ImageNet-1k under a matched recipe. (needs a GPU you connect) - **Robustness under corruptions** — Accuracy of pretrained models on CIFAR-10-C by corruption type and severity. - **Train a detector or segmenter** — Detection/segmentation model trained and evaluated on COCO-style data. (needs a GPU you connect) - **Label-noise and test-set overfitting audit** — Find mislabeled test items and compare accuracy on a re-collected test set (CIFAR-10.1). - **Novel-view synthesis benchmark** — NeRF/3D Gaussian splatting quality (PSNR/SSIM/LPIPS) on a standard scene set. (needs a GPU you connect) ### Formal methods - **Model-check a TLA+ specification** — TLC run over a bounded instance: invariants hold or a counterexample trace. - **Find models or counterexamples with Alloy** — Alloy analysis of a relational model within a stated scope. - **Symbolic execution of Python against its contracts** — Counterexamples to a Python function's pre/postconditions. - **Property-based testing with shrinking** — Minimal failing input for a stated property of real code. - **Machine-checked proof in Coq** — A theorem stated and proved in Coq using the standard library. - **Is this program transformation sound?** — SMT encoding of a transformation with a proof or counterexample. ### Human-computer interaction - **Analyse my controlled user study** — Mixed-effects analysis of a within-subjects experiment with effect sizes and CIs. - **Power analysis and preregistration** — Required sample size and a preregistration document with the analysis script dry-run on simulated data. - **Fit Fitts' law to pointing data** — Model fit, throughput and per-condition comparison from pointing trials. - **Code interview transcripts** — Codebook, coded excerpts and inter-rater agreement for qualitative data the researcher provides. - **Analyse deployment logs** — Usage patterns and retention from in-the-wild interaction logs. - **Survey instrument analysis** — Reliability (alpha/omega) and factor structure of a questionnaire. - **Reproduce a paper from its OSF artifact** — Re-run a published analysis from its OSF materials and compare numbers. ### Information retrieval - **BM25 ad-hoc evaluation on a test collection** — nDCG/MAP/recall of BM25 on a public collection with qrels. - **Dense retrieval across the full BEIR suite** — Zero-shot dense-retriever results on all BEIR datasets. (needs a GPU you connect) - **Statistically sound system comparison** — Paired tests with multiple-comparison control over per-query metrics from run files. - **Retrieval-augmented generation evaluation** — Retrieval and answer quality of a RAG pipeline on a QA benchmark. (needs a GPU you connect) - **Corpus and query-log analytics** — SQL analytics over a large document or interaction corpus. ### Machine learning - **Is an improvement real or seed noise?** — Seed-level significance test and effect size for a claimed gain. - **Fit a scaling law** — Scaling exponents with bootstrap intervals from a run table. - **Ablation study of a small model** — Causal ablation of components on a CPU-sized task with seeds and compute matched. - **Budgeted hyperparameter search on tabular benchmarks** — Tuned baselines on an OpenML suite with the search budget disclosed. - **Compare architectures at GPU scale** — Matched-recipe training runs of models too large for CPU. (needs a GPU you connect) - **Language-model benchmark evaluation** — lm-evaluation-harness scores for open models with the harness version pinned. (needs a GPU you connect) - **Benchmark contamination audit** — N-gram/near-duplicate overlap of a benchmark with a pretraining-corpus sample. - **Reproduce a published result at CPU scale** — Re-run a paper's code from GitHub and compare its numbers. ### Natural language and LLMs - **Tokenisation and perplexity experiments** — Tokenizer behaviour and perplexity of small models on a corpus. - **Score translations or summaries** — BLEU, chrF and ROUGE with signatures for system outputs. - **Fine-tune a small language model** — Fine-tuned small model (≤ ~125M params) with evaluation against the base. - **LoRA fine-tune of a billion-parameter model** — Parameter-efficient fine-tuning of a 1B+ model with a disclosed search budget. (needs a GPU you connect) - **Social-bias probing of open models** — BBQ/CrowS-Pairs-style results disaggregated by group. (needs a GPU you connect) - **Semantic search over a corpus** — Embedding index and retrieval quality over a document set. - **Text-classification baselines** — TF-IDF/linear and small-transformer baselines on a public dataset with error analysis. - **Serve a model for high-throughput inference** — Throughput/latency of a served model under load (vLLM). (needs a GPU you connect) ### Robotics and control - **Simulate a robot and close a control loop** — MuJoCo/PyBullet simulation with a feedback controller and tracking metrics. - **Classical control design** — LQR, pole placement, Bode and root-locus analysis of a plant. - **State estimation with Kalman-family filters** — EKF/UKF/particle-filter estimates and consistency checks on a dataset. - **Motion-planning benchmark** — OMPL planners compared on success rate and path quality. - **Embodied navigation in photorealistic scenes** — Habitat navigation agents on standard scene datasets. (needs a GPU you connect) - **Pose-graph SLAM back end** — Least-squares pose-graph optimisation with loop closures on a dataset. ### Security and cryptography - **Find an input that reaches a code location** — Symbolic execution of a binary to a target block (e.g. a crackme's success path). - **Analyse a binary's structure** — Sections, symbols, disassembly and emulation of an ELF binary. - **Break a textbook cryptographic construction** — Working attack (padding oracle, CBC bit-flip, small-exponent RSA) with explanation. - **Vulnerability trends from NVD** — CVE counts, CWE mix and CVSS trends for a product or class. - **Scan a sample corpus with YARA** — Rule matches and false-positive checks over files the researcher supplies. - **Membership-inference evaluation** — Attack AUC against a trained model and the effect of a defence. ### Systems and databases - **Analytical SQL on large Parquet data** — Query results, plans and engine comparison over multi-GB columnar data. - **Query-optimiser study on TPC-H** — Plan quality and regret across queries and scale factors. - **Deterministic simulation testing of a protocol** — A distributed protocol run under a seeded scheduler until a correctness bug appears, with the replay. - **Find a tail-latency regression** — Percentile analysis with CIs and change-point detection on latency measurements. - **Re-run a reproducibility package** — A paper's artifact re-executed from Zenodo/GitHub with results compared.