---------------------------------------------------------------------- This is the API documentation for the labrat library. ---------------------------------------------------------------------- ## Query models and clients Shared results and clients for external biological resources. ## Biological queries Gene, variant, and biomedical literature queries. ## Project and file management Classes for projects, archives, and scientific files. ## ProjectManager methods ## FileOrganizer methods ## Genetics DNA and protein sequence utilities. ## Laboratory math Common concentration and optical calculations. ---------------------------------------------------------------------- This is the CLI documentation for the package. ---------------------------------------------------------------------- ## CLI: labrat ``` Usage: labrat [OPTIONS] COMMAND [ARGS]... Labrat - A basic science lab framework for reproducibility and lab management. Options: --help Show this message and exit. Commands: archive Archive a directory. organize Organize files in Downloads and Documents directories. project Manage projects. query Query public gene, variant, and biomedical literature resources. ``` ### labrat query ``` Usage: labrat query [OPTIONS] COMMAND [ARGS]... Query public gene, variant, and biomedical literature resources. Options: --help Show this message and exit. Commands: gene Find gene annotations through MyGene. literature Find biomedical publications through PubTator 3. variant Find an rsID or hg19 genomic HGVS identifier through... ``` ### labrat query gene ``` Usage: labrat query gene [OPTIONS] GENE Find gene annotations through MyGene. Options: --species TEXT [default: human] --limit INTEGER RANGE [default: 5; 1<=x<=100] --all-matches Display lower-ranked matches. --format [table|json] --help Show this message and exit. ``` ### labrat query variant ``` Usage: labrat query variant [OPTIONS] VARIANT Find an rsID or hg19 genomic HGVS identifier through MyVariant. Options: --limit INTEGER RANGE [default: 5; 1<=x<=100] --format [table|json] --help Show this message and exit. ``` ### labrat query literature ``` Usage: labrat query literature [OPTIONS] [SEARCH_TEXT] Find biomedical publications through PubTator 3. Options: --gene TEXT Resolve and search for a gene concept. --disease TEXT Resolve and search for a disease concept. --variant TEXT Resolve and search for a variant concept. --chemical TEXT Resolve and search for a chemical concept. --relation [any|associate|cause|compare|convert|cotreat|drug_interact|inhibit|interact|negative_correlate|positive_correlate|prevent|stimulate|treat] Require an extracted relation. --page INTEGER RANGE [default: 1; x>=1] --limit INTEGER RANGE [default: 10; 1<=x<=100] --format [table|json] --help Show this message and exit. ``` ### labrat project ``` Usage: labrat project [OPTIONS] COMMAND [ARGS]... Manage projects. Options: --help Show this message and exit. Commands: delete Delete a project (archives it first). list List all projects. new Create a new project. ``` ### labrat project new ``` Usage: labrat project new [OPTIONS] Create a new project. Options: --type TEXT Type of project (e.g., computational-biology, data- science) [required] --name TEXT Name of the project [required] --path PATH Path where the project will be created [required] --description TEXT Description of the project [required] --username TEXT Username for project manager (defaults to system default) --help Show this message and exit. ``` ### labrat project list ``` Usage: labrat project list [OPTIONS] List all projects. Options: --username TEXT Username for project manager (defaults to system default) --help Show this message and exit. ``` ### labrat project delete ``` Usage: labrat project delete [OPTIONS] Delete a project (archives it first). Options: --path PATH Path to the project to delete [required] --archive-dir PATH Directory where the archived project will be stored [required] --username TEXT Username for project manager (defaults to system default) --yes Confirm the action without prompting. --help Show this message and exit. ``` ### labrat archive ``` Usage: labrat archive [OPTIONS] Archive a directory. Options: --source PATH Source directory to archive [required] --destination PATH Base directory for storing archives [required] --name TEXT Name for the archive [required] --help Show this message and exit. ``` ### labrat organize ``` Usage: labrat organize [OPTIONS] Organize files in Downloads and Documents directories. By default, scientific data files (fastq, fasta, sam, bam, vcf, fits, hdf5, nc, etc.) are moved to Documents/Research_Data. Use --science-dir to specify a custom location. Examples: labrat organize --science labrat organize --science --science-dir ~/Research labrat organize --keyword "project_alpha" labrat organize --all Options: --science Organize scientific data files (fastq, fasta, sam, bam, vcf, fits, hdf5, etc.) to Documents/Research_Data --science-dir PATH Custom directory for scientific data files (default: Documents/Research_Data) --keyword TEXT Move files containing this keyword to a specific folder --pictures Organize picture files to Pictures folder --videos Organize video files to Videos folder --archives Organize archive files by compression type --all Organize all file types --help Show this message and exit. ``` ---------------------------------------------------------------------- This is the User Guide documentation for the package. ---------------------------------------------------------------------- ### User Guide Labrat provides command-line and Python tools designed to improve reproducibility, simplify laboratory management, and support common biomedical research tasks. ## Start here - [Install Labrat](01-installation.qmd) and verify the command-line interface. - [Query biological databases/APIs](02-biological-queries.qmd) such as MyGene, MyVariant, and PubTator 3. - [Manage projects](03-project-management.qmd) from a reusable template and keep a local project list. - [Organize files and folders](04-files-and-archives.qmd). - [Inspect and translate sequences](05-sequence-utilities.qmd) from Python. ## Choose an interface Use the command-line interface for interactive work and shell pipelines: ```bash labrat --help labrat query --help ``` Use the Python API when a result needs to remain in memory, become part of an analysis, or feed another reproducible step: ```python from labrat.query import query_gene result = query_gene("BMPR2") print(result.provider) print(result.retrieved_at) ``` ## Query Limitations Query results depend on live external services and their current database builds. Labrat records retrieval time and provider metadata. Save JSON output to preserve the returned content used in an analysis or report. File organization commands modify local files. Review the relevant guide and the files in `Downloads` and `Documents` before running them. ## Getting Started ## Requirements Labrat requires Python 3.10 or later and `pip`. ## Install from PyPI ```bash pip install pylabrat ``` An isolated environment keeps Labrat and its dependencies separate from the system Python installation: ```bash python -m venv .venv source .venv/bin/activate python -m pip install --upgrade pip python -m pip install pylabrat ``` ## Install from source ```bash git clone https://github.com/sdhutchins/labrat.git cd labrat pip install . ``` Use an editable installation while developing Labrat: ```bash pip install -e . ``` ## Verify the installation ```bash labrat --help labrat query --help ``` Confirm that the Python package is available in the same environment: ```bash python -c "import labrat; print(labrat.__file__)" ``` ## Local state Project and file-management features use `~/.labrat/` for configuration and logs. The project manager creates `~/.labrat/config.json` when it first needs the project list. ## Build the documentation Great Docs requires Python 3.11 or later and a separate Quarto installation. Install the documentation dependencies and build the site from the repository root: ```bash pip install -e ".[docs]" great-docs build ``` The generated site is written to `great-docs/_site/`. Preview it locally with: ```bash great-docs preview ``` The source pages live in `docs/`. Great Docs regenerates `great-docs/` during each build, and the directory remains outside version control. ## Guides ### Biological Queries Labrat queries MyGene, MyVariant, and PubTator 3 while retaining provider and retrieval provenance. The default terminal view is intended for reading. Use JSON when a response will be saved, parsed, or cited in an analysis. ## Genes ```bash labrat query gene BMPR2 labrat query gene BMPR2 --all-matches ``` Gene queries default to human records. Use `--species` to select another MyGene-supported species. The default view displays the result with the highest MyGene `_score` and notes whether additional matches were returned. The score ranks records within the current search. The top panel can include the gene symbol, name, Entrez identifier, Ensembl identifier, UniProt accession, RefSeq accessions, taxonomy, and summary when MyGene reports them. Increase `--limit` to retrieve more candidates, then add `--all-matches` to display the lower-ranked records: ```bash labrat query gene ALK --limit 10 --all-matches ``` Use `--species` when a symbol is being resolved outside humans: ```bash labrat query gene Bmpr2 --species mouse ``` ## Variants ```bash labrat query variant rs429358 ``` The command accepts an rsID or an hg19 genomic HGVS identifier. An rsID is sent as a fielded dbSNP query to reduce unrelated fuzzy matches. Each returned allele gets its own panel because one identifier can map to multiple alleles or records. The readable view summarizes dbSNP identifiers, genes, ClinVar labels, CADD PHRED scores, and overall gnomAD exome and genome allele frequencies when those sources report them. MyVariant's primary genomic identifiers use the hg19 assembly. Responses may include mapped coordinates for other assemblies when a source provides them. ClinVar labels are condition-specific and should be interpreted with their associated records. ## Literature Search PubTator 3 with free text: ```bash labrat query literature "BMPR2 pulmonary arterial hypertension" ``` Resolve normalized biological concepts before searching: ```bash labrat query literature --gene BMPR2 \ --disease "pulmonary arterial hypertension" ``` Require a text-mined relation between exactly two concepts: ```bash labrat query literature --gene BMPR2 \ --disease "pulmonary arterial hypertension" \ --relation associate ``` PubTator relations are text-mined associations. They should be interpreted with the cited publications. Structured options can resolve genes, diseases, variants, and chemicals: ```bash labrat query literature --variant rs429358 --disease "Alzheimer disease" labrat query literature --chemical sildenafil --disease \ "pulmonary arterial hypertension" ``` Use either free text or structured options in one command. A relation query requires exactly two structured entities. Labrat also requires an exact unique autocomplete name before constructing a semantic query. An ambiguous concept returns candidate names so you can choose the exact match. `--page` selects a PubTator result page. `--limit` controls how many records are shown in the terminal, but Labrat retains the complete returned page in JSON. ## Machine-readable output All query commands accept `--format json`. JSON output retains the complete provider response and query provenance for downstream analysis. ```bash labrat query gene BMPR2 --format json > bmpr2-mygene.json labrat query variant rs429358 --format json > rs429358-myvariant.json ``` Each serialized `QueryResult` contains: - `kind`, the Labrat query type; - `query`, the provider query or normalized semantic query; - `provider`, the external resource name; - `retrieved_at`, an ISO 8601 UTC timestamp; - `metadata`, including available build and source information; - `data`, the complete provider response retained by Labrat. ## Python API The Python API returns the same immutable `QueryResult` model used by the CLI: ```python from labrat.query import query_gene result = query_gene("BMPR2", species="human", limit=5) highest_provider_hit = result.data["hits"][0] print(result.provider) print(result.metadata["build_version"]) print(highest_provider_hit["symbol"]) ``` MyGene normally returns hits in provider rank order. The Rich terminal renderer explicitly sorts them by `_score` before choosing the displayed top match. Code that selects a record should use identifiers and taxonomy to confirm identity, with list position serving only as provider rank. ## Expected failures Labrat converts provider errors, timeouts, unexpected response shapes, and ambiguous PubTator concepts into readable query errors. Empty results are valid responses and appear as an empty-result panel. ### Project Management Labrat can create projects from Cookiecutter templates and record their paths in a local registry. This is useful when several analyses need the same initial directory structure and metadata. ## Create a project Choose a template, a display name, a parent directory, and a description: ```bash labrat project new \ --type computational-biology \ --name "BMPR2 Variant Analysis" \ --path ./projects \ --description "Analyze BMPR2 variants in pulmonary hypertension" ``` The default template names are `computational-biology`, `data-science`, and `andymcdgeo-streamlit`. Templates are remote repositories, so first-time project creation requires network access. `--path` identifies the parent directory. Cookiecutter creates the project inside that directory. Labrat replaces spaces in the submitted project name with underscores before passing it to the template. The template can still control the final directory name. ## Inspect your projects ```bash labrat project list ``` Labrat stores the absolute project path, description, project type, creation time, and last-modified time in `~/.labrat/config.json`. The project list records pointers to the original project directories. Keep project files, version-control history, and separate backups as the authoritative records of an analysis. ## Use the Python API ```python from pathlib import Path from labrat.project import ProjectManager manager = ProjectManager(username="Dr. Jane Doe") projects = manager.list_projects() for project in projects: print(project["name"], Path(project["path"])) ``` The Python API also exposes methods to update timestamps, recursively list project files, add templates, and archive projects. See the generated `ProjectManager` API reference for their full signatures. ## Delete with an archive ```bash labrat project delete \ --path ./projects/BMPR2_Analysis \ --archive-dir ./project-archives ``` This command asks for confirmation. It copies and zips the project before recursively deleting the original directory and removing its registry entry. The archive parent directory must already exist. ::: {.callout-warning} Confirm that the archive was written to the intended filesystem before relying on it for recovery. Deleting a project removes the original directory tree. ::: To preserve the source directory while creating a backup, use `labrat archive`. ### File and Folder Organization Labrat provides two distinct file workflows. Archiving copies a directory and creates a ZIP file. Organizing moves matching files out of the top level of `Downloads` and `Documents`. ## Create an archive Create the destination parent first, then archive a source directory: ```bash mkdir -p ./archives labrat archive \ --source ./my_project \ --destination ./archives \ --name my_project ``` Labrat creates a timestamped directory such as `my_project_archive_20260905_143000` and a ZIP archive beside it. The original source remains in place. Archive names include local time down to the second. Run archives sequentially when they share a project name and destination. ## Organize scientific files ```bash labrat organize --science ``` This command moves recognized scientific files from the top level of `~/Downloads` and `~/Documents` into `~/Documents/Research_Data`. Recognized formats include common sequence, alignment, variant, annotation, phylogenetic, structure, astronomy, HDF5, and NetCDF extensions. Choose a different destination with: ```bash labrat organize --science --science-dir ./research-data ``` `--science-dir` selects the destination. Labrat continues scanning the top level of `~/Downloads` and `~/Documents`. ## Organize by filename or file category ```bash labrat organize --keyword project_alpha labrat organize --archives labrat organize --pictures --videos ``` Keyword matching is case-sensitive and moves matching files into `~/Documents/Organized_Files`. Archive files are grouped by compression type under `~/Documents/Archive`. The current media implementation uses one shared organizer. Supplying either `--pictures` or `--videos` processes both recognized picture and video files. ## Review before running ::: {.callout-warning} The organize commands move files immediately. Review the top level of `~/Downloads` and `~/Documents` before running them. ::: If a destination already contains a file with the same name, Labrat compares modification times. It deletes the older copy and keeps the newer copy. Equal timestamps cause the source file to be deleted. Review duplicate filenames before organizing irreplaceable data. `labrat organize --all` runs every organizer, including a keyword-specific file move. For predictable behavior, use the narrowest explicit option for the files you intend to move. File operations are recorded under `~/.labrat/`. ### Sequence Utilities Labrat provides focused Python utilities for nucleotide composition, DNA complements, and FASTA translation. They fit quick checks in scripts, notebooks, and teaching examples. ## Count canonical nucleotides ```python from labrat.genetics import atgc_content counts = atgc_content("ATGCGAT") print(counts) ``` The result is: ```text {'A': 2, 'T': 2, 'G': 2, 'C': 1} ``` Input is case-insensitive. Counts include `A`, `T`, `G`, and `C` only, so the total can be smaller than the submitted sequence length. ## Create a complementary strand ```python from labrat.genetics import complementary_dna complement = complementary_dna("ATGN") print(complement) ``` The result is `TACX`. Non-canonical characters become `X`. The function returns the direct complement in the submitted order. Reverse this sequence when you need the reverse complement. ## Translate a FASTA sequence Create a small FASTA file whose nucleotide count is divisible by three: ```text >example ATGAAATGG ``` Translate it with: ```python from labrat.genetics import dna2aminoacid protein = dna2aminoacid("example.fasta") print(protein) ``` The result is `MKW`. Labrat removes whitespace, accepts lowercase sequence characters, and skips every line beginning with `>`. Stop codons are returned as `_`, and translation continues with the next codon. Translation raises an error for an empty sequence, a length that leaves a remainder after division by three, or a codon containing a character outside `A`, `T`, `G`, and `C`. ## Laboratory calculations Labrat also provides laboratory calculations through its Python API. These functions accept plain numeric values, so confirm the intended units before using a result in experimental work.