awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
biopython avatar

biopython/biopython

0
View on GitHub↗
5,078 stars·1,913 forks·Python·17 viewsbiopython.org↗

Biopython

Biopython is a bioinformatics library for Python providing tools to parse, manipulate, and analyze biological sequences, molecular structures, and phylogenetic trees. It serves as a biological sequence parser for genomic and proteomic data across multiple industry-standard file formats and acts as an interface for querying biological data and citations from NCBI Entrez repositories.

The project distinguishes itself through specialized toolkits for protein structure analysis and phylogenetic tree construction. It includes a protein structure analyzer for processing PDB and mmCIF files to calculate molecular geometry, as well as a phylogenetic tree toolkit for analyzing evolutionary relationships between species.

The library covers a broad range of bioinformatics capabilities, including genomic sequence analysis for transcription and translation, the management of sequence alignments, and population genetic calculations. It also provides structural analysis tools for 3D atomic coordinate manipulation, as well as utilities for genomic feature visualization and biogeographical data modeling.

The system integrates with external bioinformatics binaries via tool wrapping and supports persistent biological record storage through SQL-backed sequence storage.

Features

  • Biological Format Parsers - Uses a consistent interface to translate diverse biological file formats into a standardized internal record object.
  • Biological Sequence Objects - Handles biological sequences using specialized objects that provide domain-specific methods beyond basic strings.
  • Biological Sequence Parsers - Provides a comprehensive system for reading and writing genomic and proteomic data across multiple industry-standard file formats.
  • Bioinformatics Libraries - Provides a comprehensive collection of computational tools for processing and analyzing biological data and genomic sequences.
  • Genomic Data Analysis - Parses and manipulates DNA and RNA sequences to identify features and generate reverse complements.
  • Protein Databases - Provides automated requests to retrieve protein structures and datasets from the Protein Data Bank.
  • Sequence Alignment - Implements pairwise and multiple sequence alignment algorithms to identify similarities between biological sequences.
  • Biology and Bioinformatics - Provides a comprehensive set of computational tools for genomic and biological data analysis.
  • Biological Database Interfaces - Interfaces with relational databases to store and query large-scale biological datasets.
  • Biological Sequence Parsers - Reads and writes biological sequence data in various formats using a unified interface.
  • Sequence Analysis - Removes noise and unwanted artifacts from raw biological sequence data to improve analysis quality.
  • Database Clients - Implements a client for connecting to and querying NCBI Entrez repositories to fetch biological data and citations.
  • Molecular Structure Parsing - Converts PDB and mmCIF files into hierarchical structure objects or direct dictionary maps.
  • Sequence Alignment File Loaders - Parses sequence alignment files into a structured internal representation for analysis.
  • Biological Record Management - Stores biological sequences with their identifiers, names, descriptions, and associated metadata.
  • Wildlife and Biology APIs - Provides programmatic interfaces to retrieve genomic and proteomic data from various biological repositories.
  • Evolutionary Tree Visualizations - Processes and visualizes evolutionary trees to determine relationships between biological species.
  • Phylogenetic Tree Construction - Builds and analyzes evolutionary trees to study species relationships based on genetic distance and biogeographical data.
  • Biological Transcriptions - Converts DNA sequences to RNA or RNA sequences back to DNA by substituting thymine and uracil.
  • Genetic Code Translations - Converts nucleotide sequences into amino acids using genetic code tables tailored to different organisms.
  • Genetic Translation Tools - Converts DNA or RNA sequences into protein sequences using genetic code translation tables.
  • Molecular Structural Hierarchies - Organizes 3D molecular data into a nested tree of structures, models, chains, residues, and atoms.
  • Nucleotide Sequence Manipulations - Creates new biological sequences by replacing each base with its complement and reversing the order.
  • Phylogenetic Tree Toolkits - Provides functions for constructing and analyzing evolutionary trees to study relationships between biological species.
  • Protein Structure Analysis - Processes 3D atomic coordinates from PDB files to measure molecular geometry and analyze residue exposure.
  • Geometry Analyzers - Processes PDB and mmCIF files to calculate molecular geometry and residue solvent exposure.
  • PDB File Parsing - Reads Protein Data Bank files and removes disordered atoms to clean structural data.
  • Vector Geometry Utilities - Uses vector geometry to calculate molecular distances and angles based on 3D atomic coordinates.
  • Protein Sequence Translation - Translates gene sequences into predicted protein products using genomic feature files.
  • XML Parsing - Parses XML data retrieved from remote biological repositories via API requests.
  • Genome Annotation - Extracts structural annotations and coordinates from GFF files to map genomic features.
  • Metadata Retrieval - Fetches biological metadata for gene identifiers from external genomic databases to annotate sequences.
  • Genome Visualization - Creates biological sequence diagrams and genomic schematics in various vector and bitmap formats.
  • Biological Sequence Export - Converts biological sequence records into formatted strings compatible with industry-standard bioinformatics output.
  • Biological Sequence Writing - Exports sequence record objects to files or strings using supported bioinformatics formats.
  • Non-Coding Region Extraction - Isolates non-coding DNA sequences located between genes from a larger genomic sequence.
  • Record-Level Indexes - Creates mappings of sequence identifiers to file locations to enable fast random access within large datasets.
  • Molecular Structure Export - Writes structure objects to PDB files while filtering specific models, chains, residues, or atoms.
  • Molecular Property Calculation - Computes molecular weight, masses, and composition complexity for biological sequences.
  • Database Search Integration - Processes and manages search results retrieved from external bioinformatics databases via integrated query interfaces.
  • SQL-Backed Stores - Provides SQL-backed storage to persist and partition biological sequence records using named namespaces.
  • Random Access Data Retrieval - Implements index-based random access to retrieve specific biological sequences from large files using offsets.
  • Molecular Structural Navigation - Traverses structural data using a hierarchical model of structures, models, chains, residues, and atoms.
  • Sequence Metadata Management - Attaches dictionary-based annotations and per-letter metadata, such as quality scores, to biological sequences.
  • Sequence File Indexing - Creates disk or memory based indexes for large sequence files to enable fast random access.
  • Format Conversions - Transforms biological sequence data between different file formats to ensure interoperability across bioinformatics tools.
  • Bioinformatics Tool Adapters - Wraps external command-line bioinformatics binaries to execute complex analysis pipelines within Python.
  • Bioinformatics Tool Wrapping - Automates analysis pipelines by wrapping command-line bioinformatics tools for use within scripts.
  • Bioinformatics Binary Wrapping - Connects to phylogenetic analysis tools like PAML to perform evolutionary modeling based on maximum likelihood.
  • 3D Geometry Measurers - Calculates distances, bond angles, and torsion angles using vector representations of atomic coordinates.
  • Atomic Contact Detectors - Detects neighboring atoms within a specified distance using a fast spatial search tree.
  • Atomic Coordinate Manipulators - Performs 3D vector operations and coordinate transformations to update atom positions.
  • Genetic Coordinate Transformations - Translates positions between different genetic coordinate systems or reference frames.
  • Atomic Disorder Management - Manages disordered atoms and point mutations by encapsulating alternate positions within single objects.
  • Polypeptide Sequence Extractors - Builds polypeptide objects from structure data to derive amino acid sequences and atom lists.
  • Protein Structural Refinement - Manipulates structural biology data by adding hydrogens, identifying disulfide bridges, and renumbering residues.
  • RNA Structure Processors - Parses RNA secondary structures and manages sequence objects representing modified nucleotides.
  • Structural Bioinformatics Analysis - Determines secondary structure and accessible surface area by integrating with external structural programs.
  • Structural Model Alignments - Maps residues between related protein structures and calculates matrices to minimize distance between atom sets.
  • Scientific Computing - Tools for biological computation.
  • Bioinformatics Libraries - Python-based biological computing tools and documentation.
  • Scientific Computing - Listed in the “Scientific Computing” section of the Awesome Python awesome list.
  • Code Libraries - Listed in the “Code Libraries” section of the Awesome Bioie awesome list.

Star history

Star history chart for biopython/biopythonStar history chart for biopython/biopython

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does biopython/biopython do?

Biopython is a bioinformatics library for Python providing tools to parse, manipulate, and analyze biological sequences, molecular structures, and phylogenetic trees. It serves as a biological sequence parser for genomic and proteomic data across multiple industry-standard file formats and acts as an interface for querying biological data and citations from NCBI Entrez repositories.

What are the main features of biopython/biopython?

The main features of biopython/biopython are: Biological Format Parsers, Biological Sequence Objects, Biological Sequence Parsers, Bioinformatics Libraries, Genomic Data Analysis, Protein Databases, Sequence Alignment, Biology and Bioinformatics.

Which projects share features with biopython/biopython?

Projects with overlapping indexed features include: attractivechaos/klib — klib is a comprehensive C standard library extension and data structure toolkit. It provides a set of fundamental… k-dense-ai/claude-scientific-skills — This project is a scientific agent framework and workflow orchestrator designed to extend large language models with… arcinstitute/evo2 — evo2 is a genomic large language model and foundation model designed to predict, generate, and analyze genetic… bdvajstudio/javdb — javdb is a mobile adult media indexer and video database browser. It functions as a REST API client that allows users… google-deepmind/alphafold3 — AlphaFold3 is a biomolecular structure prediction model and bioinformatics structural analysis tool. It uses a deep… spaceandtimefdn/sxt-python-sdk — The SxT-Python-SDK is a Python library and SQL database client designed for executing queries and managing database…

Projects sharing features with Biopython

These projects share indexed features with Biopython. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • attractivechaos/klibattractivechaos avatar

    attractivechaos/klib

    4,679View on GitHub↗

    klib is a comprehensive C standard library extension and data structure toolkit. It provides a set of fundamental tools for memory management, data organization, and general-purpose utility functions for standalone C applications. The project features specialized capabilities for bioinformatics sequence analysis, including the parsing of FASTA, FASTQ, and Newick formats and the implementation of Smith-Waterman sequence alignment and Hidden Markov Models. It also includes a mathematical computation library for numerical routines and expression evaluation, as well as a lightweight HTTP and FTP

    C
    View on GitHub↗4,679
  • bdvajstudio/javdbbdvajstudio avatar

    bdvajstudio/javdb

    3,594View on GitHub↗

    javdb is a mobile adult media indexer and video database browser. It functions as a REST API client that allows users to search for and discover adult video metadata and titles from a remote database. The application includes a magnet link search tool that automatically locates downloadable torrent files associated with specific media entries by querying external search engines. The system manages content discovery through video database browsing and the retrieval of media information via standard HTTP requests.

    javbusjavdbjavhd
    View on GitHub↗3,594
  • arcinstitute/evo2ArcInstitute avatar

    ArcInstitute/evo2

    3,951View on GitHub↗

    evo2 is a genomic large language model and foundation model designed to predict, generate, and analyze genetic information across different species. It functions as a nucleotide sequence modeler and a DNA sequence generator, using transformer-based sequence modeling to process genomic data. The system provides capabilities for synthetic DNA generation, creating new genetic sequences based on biological prompts or species-specific tags. It also performs nucleotide likelihood prediction to score genomic variants and analyze biological properties within DNA sequences. The model supports genomic

    Jupyter Notebook
    View on GitHub↗3,951
  • google-deepmind/alphafold3google-deepmind avatar

    google-deepmind/alphafold3

    7,613View on GitHub↗

    AlphaFold3 is a biomolecular structure prediction model and bioinformatics structural analysis tool. It uses a deep learning system to predict the three-dimensional shapes of proteins, DNA, RNA, and ligands. The system functions as a diffusion-based protein folding model that predicts the spatial coordinates of biomolecular atoms and interactions. It utilizes a GPU-accelerated inference pipeline to process genetic sequences and structural templates for molecular modeling. The project covers structural bioinformatics analysis and protein interaction modeling to determine the physical arrangem

    Python
    View on GitHub↗7,613
Compare all 30 related projects→