awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

10 dépôts

Awesome GitHub RepositoriesSequence Analysis

Tools for analyzing ordered data sequences like biological or time-series data.

Explore 10 awesome GitHub repositories matching data & databases · Sequence Analysis. Refine with filters or upvote what's useful.

Awesome Sequence Analysis GitHub Repositories

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • josephmisiti/awesome-machine-learningAvatar de josephmisiti

    josephmisiti/awesome-machine-learning

    72,867Voir sur GitHub↗

    This project is a comprehensive, community-driven directory of machine learning resources, software libraries, and educational materials. It serves as a centralized knowledge base for developers and researchers, organizing tools and frameworks by their primary programming language and technical domain to simplify discovery across the artificial intelligence ecosystem. The collection distinguishes itself by providing a cross-language development index that spans diverse programming environments, including C, C++, Rust, Clojure, and Python. It covers a wide range of specialized capabilities, fr

    Integrates probabilistic models for analyzing ordered data sequences and time-series information.

    Python
    Voir sur GitHub↗72,867
  • google-research/google-researchAvatar de google-research

    google-research/google-research

    38,139Voir sur GitHub↗

    This repository serves as a comprehensive research platform and toolkit for advancing machine learning, quantum computing, and large-scale scientific data analysis. It provides foundational frameworks for developing complex algorithmic systems, offering the necessary infrastructure for distributed training, computational graph execution, and high-performance model development. The project distinguishes itself by integrating specialized research domains with robust, privacy-preserving methodologies. It supports diverse scientific discovery through tools for quantum simulation, physics-informed

    Applies deep learning to process and interpret genetic data for variant identification.

    Jupyter Notebookaimachine-learningresearch
    Voir sur GitHub↗38,139
  • mission-peace/interviewAvatar de mission-peace

    mission-peace/interview

    11,306Voir sur GitHub↗

    This project is a comprehensive library of reference implementations for fundamental data structures and algorithms, designed to support technical interview preparation and software engineering assessments. It provides a structured collection of computational techniques for solving complex problems involving arrays, strings, graphs, trees, and mathematical analysis. The library distinguishes itself by offering specialized implementations for advanced topics, including concurrent programming patterns and geometric algorithms. It features thread-safe primitives for managing shared state and tas

    Provides algorithms for identifying the longest chain of consecutive integers within two-dimensional arrays.

    Java
    Voir sur GitHub↗11,306
  • k-dense-ai/claude-scientific-skillsAvatar de K-Dense-AI

    K-Dense-AI/claude-scientific-skills

    8,907Voir sur GitHub↗

    This project is a scientific agent framework and workflow orchestrator designed to extend large language models with specialized tools for genomic, chemical, and biological research. It provides a system for planning research hypotheses and executing automated workflows by integrating scientific databases with dynamic code execution. The framework includes a cheminformatics modeling suite for predicting molecular bioactivity and performing virtual screening, alongside a bioinformatics analysis toolkit for processing genomic sequences and single-cell data. It also features an academic document

    Analyzes DNA and protein sequences to annotate genetic variants and identify pathogenicity.

    Pythonai-scientistbioinformaticschemoinformatics
    Voir sur GitHub↗8,907
  • deepchem/deepchemAvatar de deepchem

    deepchem/deepchem

    6,545Voir sur GitHub↗

    DeepChem is an open-source Python framework for applying deep learning to molecular, chemical, and biological data, serving as a comprehensive toolkit for drug discovery and materials science. At its core, it provides a featurizer-pipeline abstraction that converts raw molecular data into numerical representations, including graph-based molecular structures, SMILES tokenization vocabularies, and disk-sharded dataset persistence for handling large-scale data that exceeds RAM capacity. The framework distinguishes itself through integrated molecular docking workflows that automate pocket detecti

    Featurizes genomic and proteomic sequences from alignment files for downstream machine learning models.

    Pythonbiologydeep-learningdrug-discovery
    Voir sur GitHub↗6,545
  • biopython/biopythonAvatar de biopython

    biopython/biopython

    5,078Voir sur GitHub↗

    Biopython est une bibliothèque de bioinformatique pour Python fournissant des outils pour analyser, manipuler et étudier les séquences biologiques, les structures moléculaires et les arbres phylogénétiques. Elle sert d'analyseur de séquences biologiques pour les données génomiques et protéomiques à travers de multiples formats de fichiers standards de l'industrie et agit comme une interface pour interroger les données biologiques et les citations des dépôts NCBI Entrez. Le projet se distingue par des toolkits spécialisés pour l'analyse de structure protéique et la construction d'arbres phylogénétiques. Il inclut un analyseur de structure protéique pour traiter les fichiers PDB et mmCIF afin de calculer la géométrie moléculaire, ainsi qu'un toolkit d'arbres phylogénétiques pour analyser les relations évolutives entre les espèces. La bibliothèque couvre un large éventail de capacités bioinformatiques, incluant l'analyse de séquences génomiques pour la transcription et la traduction, la gestion des alignements de séquences et les calculs de génétique des populations. Elle fournit également des outils d'analyse structurelle pour la manipulation de coordonnées atomiques 3D, ainsi que des utilitaires pour la visualisation de caractéristiques génomiques et la modélisation de données biogéographiques. Le système s'intègre avec des binaires bioinformatiques externes via l'encapsulation d'outils et prend en charge le stockage persistant d'enregistrements biologiques via un stockage de séquences basé sur SQL.

    Removes noise and unwanted artifacts from raw biological sequence data to improve analysis quality.

    Pythonbioinformaticsbiopythondna
    Voir sur GitHub↗5,078
  • attractivechaos/klibAvatar de attractivechaos

    attractivechaos/klib

    4,679Voir sur GitHub↗

    klib est une extension complète de bibliothèque standard C et une boîte à outils de structure de données. Elle fournit un ensemble d'outils fondamentaux pour la gestion de la mémoire, l'organisation des données et des fonctions utilitaires à usage général pour les applications C autonomes. Le projet propose des capacités spécialisées pour l'analyse de séquences bioinformatiques, y compris l'analyse des formats FASTA, FASTQ et Newick et l'implémentation de l'alignement de séquences Smith-Waterman et des modèles de Markov cachés. Il inclut également une bibliothèque de calcul mathématique pour les routines numériques et l'évaluation d'expressions, ainsi qu'un client HTTP et FTP léger pour la récupération de données distantes à accès aléatoire. La boîte à outils couvre une large surface de primitives de calcul haute performance, y compris les modèles multi-threadés, la construction de tableaux de suffixes en temps linéaire et des algorithmes de tri optimisés. Elle implémente une variété de structures d'indexation de données efficaces telles que des tables de hachage avec adressage ouvert, des arbres B et des arbres AVL intrusifs, pris en charge par la gestion de séquences basée sur des pools de mémoire. Les utilitaires supplémentaires incluent l'analyse de données JSON et l'interprétation des arguments de ligne de commande.

    Provides tools for analyzing biological sequences, including FASTA/FASTQ parsing and Smith-Waterman alignment.

    C
    Voir sur GitHub↗4,679
  • jwohlwend/boltzAvatar de jwohlwend

    jwohlwend/boltz

    4,038Voir sur GitHub↗

    Boltz is a deep learning molecular modeler and biomolecular structure prediction system. It uses neural network architectures to simulate the physical folding and docking of biomolecules, specifically predicting the three-dimensional shapes of protein and ligand complexes. The project functions as a protein-ligand complex predictor and binding affinity predictor, estimating the strength and probability of molecular interactions between ligands and targets. These capabilities are applied to computer aided drug design, including ligand binding affinity prediction and protein-ligand interaction

    Transforms raw protein sequences into high-dimensional feature vectors using evolutionary information from related sequences.

    Python
    Voir sur GitHub↗4,038
  • arcinstitute/evo2Avatar de ArcInstitute

    ArcInstitute/evo2

    3,951Voir sur GitHub↗

    evo2 est un modèle de langage étendu (LLM) et un modèle de fondation génomique conçu pour prédire, générer et analyser des informations génétiques à travers différentes espèces. Il fonctionne comme un modeleur de séquences nucléotidiques et un générateur de séquences d'ADN, utilisant la modélisation de séquences basée sur les transformers pour traiter les données génomiques. Le système offre des capacités de génération d'ADN synthétique, créant de nouvelles séquences génétiques basées sur des prompts biologiques ou des tags spécifiques à une espèce. Il effectue également la prédiction de probabilité nucléotidique pour noter les variants génomiques et analyser les propriétés biologiques au sein des séquences d'ADN. Le modèle prend en charge l'analyse de séquences génomiques grâce à l'extraction de représentations de haute dimension à partir des couches intermédiaires. Ces embeddings permettent une classification spécialisée et une analyse en aval des données génétiques.

    Predicts nucleotide likelihoods across a sequence to score genomic variants or analyze biological properties.

    Jupyter Notebook
    Voir sur GitHub↗3,951
  • space-wizards/space-station-14Avatar de space-wizards

    space-wizards/space-station-14

    3,523Voir sur GitHub↗

    Space Station 14 is a C# multiplayer game and roleplay simulation framework. It is built upon an Entity-Component-System (ECS) game engine that separates logic into systems and data into components to manage complex entity interactions. The project functions as a grid-based physics simulator with a YAML data-driven prototype system for defining game objects. The project features a specialized 2D sprite rendering engine that maps server-side appearance data to client-side shaders. It implements a networking model with client-side prediction and dirty-flagged state synchronization to reduce inp

    Allows editing base pairs within a plant's genetic sequence via a hex-style interface.

    C#c-sharpgamehacktoberfest
    Voir sur GitHub↗3,523
  1. Home
  2. Data & Databases
  3. Data Analysis & Visualization
  4. Analytical Platforms and Engines
  5. Sequence Analysis

Explorer les sous-tags

  • Genomic Sequence Interpreters5 sous-tagsDeep learning models for processing and interpreting genetic data to identify variants. **Distinct from Sequence Analysis:** Distinct from general sequence analysis: focuses on genomic-specific deep learning interpretation.