7 repositorios
Deep learning models for processing and interpreting genetic data to identify variants.
Distinct from Sequence Analysis: Distinct from general sequence analysis: focuses on genomic-specific deep learning interpretation.
Explore 7 awesome GitHub repositories matching data & databases · Genomic Sequence Interpreters. Refine with filters or upvote what's useful.
This repository serves as a comprehensive research platform and toolkit for advancing machine learning, quantum computing, and large-scale scientific data analysis. It provides foundational frameworks for developing complex algorithmic systems, offering the necessary infrastructure for distributed training, computational graph execution, and high-performance model development. The project distinguishes itself by integrating specialized research domains with robust, privacy-preserving methodologies. It supports diverse scientific discovery through tools for quantum simulation, physics-informed
Applies deep learning to process and interpret genetic data for variant identification.
This project is a scientific agent framework and workflow orchestrator designed to extend large language models with specialized tools for genomic, chemical, and biological research. It provides a system for planning research hypotheses and executing automated workflows by integrating scientific databases with dynamic code execution. The framework includes a cheminformatics modeling suite for predicting molecular bioactivity and performing virtual screening, alongside a bioinformatics analysis toolkit for processing genomic sequences and single-cell data. It also features an academic document
Analyzes DNA and protein sequences to annotate genetic variants and identify pathogenicity.
DeepChem is an open-source Python framework for applying deep learning to molecular, chemical, and biological data, serving as a comprehensive toolkit for drug discovery and materials science. At its core, it provides a featurizer-pipeline abstraction that converts raw molecular data into numerical representations, including graph-based molecular structures, SMILES tokenization vocabularies, and disk-sharded dataset persistence for handling large-scale data that exceeds RAM capacity. The framework distinguishes itself through integrated molecular docking workflows that automate pocket detecti
Featurizes genomic and proteomic sequences from alignment files for downstream machine learning models.
Biopython es una biblioteca de bioinformática para Python que proporciona herramientas para analizar, manipular y estudiar secuencias biológicas, estructuras moleculares y árboles filogenéticos. Sirve como un analizador de secuencias biológicas para datos genómicos y proteómicos en múltiples formatos de archivo estándar de la industria y actúa como interfaz para consultar datos biológicos y citas de los repositorios NCBI Entrez. El proyecto se distingue por kits de herramientas especializados para el análisis de estructuras proteicas y la construcción de árboles filogenéticos. Incluye un analizador de estructuras de proteínas para procesar archivos PDB y mmCIF para calcular la geometría molecular, así como un kit de herramientas de árboles filogenéticos para analizar relaciones evolutivas entre especies. La biblioteca cubre una amplia gama de capacidades bioinformáticas, incluyendo análisis de secuencias genómicas para transcripción y traducción, gestión de alineamientos de secuencias y cálculos de genética de poblaciones. También proporciona herramientas de análisis estructural para la manipulación de coordenadas atómicas en 3D, así como utilidades para la visualización de características genómicas y modelado de datos biogeográficos. El sistema se integra con binarios bioinformáticos externos mediante envoltorios de herramientas y admite el almacenamiento persistente de registros biológicos a través de almacenamiento de secuencias respaldado por SQL.
Isolates non-coding DNA sequences located between genes from a larger genomic sequence.
Boltz is a deep learning molecular modeler and biomolecular structure prediction system. It uses neural network architectures to simulate the physical folding and docking of biomolecules, specifically predicting the three-dimensional shapes of protein and ligand complexes. The project functions as a protein-ligand complex predictor and binding affinity predictor, estimating the strength and probability of molecular interactions between ligands and targets. These capabilities are applied to computer aided drug design, including ligand binding affinity prediction and protein-ligand interaction
Transforms raw protein sequences into high-dimensional feature vectors using evolutionary information from related sequences.
evo2 is a genomic large language model and foundation model designed to predict, generate, and analyze genetic information across different species. It functions as a nucleotide sequence modeler and a DNA sequence generator, using transformer-based sequence modeling to process genomic data. The system provides capabilities for synthetic DNA generation, creating new genetic sequences based on biological prompts or species-specific tags. It also performs nucleotide likelihood prediction to score genomic variants and analyze biological properties within DNA sequences. The model supports genomic
Predicts nucleotide likelihoods across a sequence to score genomic variants or analyze biological properties.
Space Station 14 is a C# multiplayer game and roleplay simulation framework. It is built upon an Entity-Component-System (ECS) game engine that separates logic into systems and data into components to manage complex entity interactions. The project functions as a grid-based physics simulator with a YAML data-driven prototype system for defining game objects. The project features a specialized 2D sprite rendering engine that maps server-side appearance data to client-side shaders. It implements a networking model with client-side prediction and dirty-flagged state synchronization to reduce inp
Allows editing base pairs within a plant's genetic sequence via a hex-style interface.