10 रिपॉजिटरी
Tools for analyzing ordered data sequences like biological or time-series data.
Explore 10 awesome GitHub repositories matching data & databases · Sequence Analysis. Refine with filters or upvote what's useful.
This project is a comprehensive, community-driven directory of machine learning resources, software libraries, and educational materials. It serves as a centralized knowledge base for developers and researchers, organizing tools and frameworks by their primary programming language and technical domain to simplify discovery across the artificial intelligence ecosystem. The collection distinguishes itself by providing a cross-language development index that spans diverse programming environments, including C, C++, Rust, Clojure, and Python. It covers a wide range of specialized capabilities, fr
Integrates probabilistic models for analyzing ordered data sequences and time-series information.
This repository serves as a comprehensive research platform and toolkit for advancing machine learning, quantum computing, and large-scale scientific data analysis. It provides foundational frameworks for developing complex algorithmic systems, offering the necessary infrastructure for distributed training, computational graph execution, and high-performance model development. The project distinguishes itself by integrating specialized research domains with robust, privacy-preserving methodologies. It supports diverse scientific discovery through tools for quantum simulation, physics-informed
Applies deep learning to process and interpret genetic data for variant identification.
This project is a comprehensive library of reference implementations for fundamental data structures and algorithms, designed to support technical interview preparation and software engineering assessments. It provides a structured collection of computational techniques for solving complex problems involving arrays, strings, graphs, trees, and mathematical analysis. The library distinguishes itself by offering specialized implementations for advanced topics, including concurrent programming patterns and geometric algorithms. It features thread-safe primitives for managing shared state and tas
Provides algorithms for identifying the longest chain of consecutive integers within two-dimensional arrays.
This project is a scientific agent framework and workflow orchestrator designed to extend large language models with specialized tools for genomic, chemical, and biological research. It provides a system for planning research hypotheses and executing automated workflows by integrating scientific databases with dynamic code execution. The framework includes a cheminformatics modeling suite for predicting molecular bioactivity and performing virtual screening, alongside a bioinformatics analysis toolkit for processing genomic sequences and single-cell data. It also features an academic document
Analyzes DNA and protein sequences to annotate genetic variants and identify pathogenicity.
DeepChem is an open-source Python framework for applying deep learning to molecular, chemical, and biological data, serving as a comprehensive toolkit for drug discovery and materials science. At its core, it provides a featurizer-pipeline abstraction that converts raw molecular data into numerical representations, including graph-based molecular structures, SMILES tokenization vocabularies, and disk-sharded dataset persistence for handling large-scale data that exceeds RAM capacity. The framework distinguishes itself through integrated molecular docking workflows that automate pocket detecti
Featurizes genomic and proteomic sequences from alignment files for downstream machine learning models.
Biopython पायथन के लिए एक बायोइनफॉरमैटिक्स लाइब्रेरी है जो जैविक अनुक्रमों, आणविक संरचनाओं और फाइलोजेनेटिक पेड़ों को पार्स, मैनिपुलेट और विश्लेषण करने के लिए टूल्स प्रदान करती है। यह कई उद्योग-मानक फ़ाइल फॉर्मेट्स में जीनोमिक और प्रोटिओमिक डेटा के लिए एक जैविक अनुक्रम पार्सर के रूप में कार्य करती है और NCBI Entrez रिपॉजिटरी से जैविक डेटा और उद्धरणों को क्वेरी करने के लिए एक इंटरफेस के रूप में कार्य करती है। यह प्रोजेक्ट प्रोटीन संरचना विश्लेषण और फाइलोजेनेटिक ट्री निर्माण के लिए विशेष टूलकिट्स के माध्यम से खुद को अलग करता है। इसमें आणविक ज्यामिति की गणना करने के लिए PDB और mmCIF फ़ाइलों को प्रोसेस करने के लिए एक प्रोटीन संरचना एनालाइज़र, और प्रजातियों के बीच विकासवादी संबंधों का विश्लेषण करने के लिए एक फाइलोजेनेटिक ट्री टूलकिट शामिल है। यह लाइब्रेरी बायोइनफॉरमैटिक्स क्षमताओं की एक विस्तृत श्रृंखला को कवर करती है, जिसमें ट्रांसक्रिप्शन और अनुवाद के लिए जीनोमिक अनुक्रम विश्लेषण, अनुक्रम संरेखण का प्रबंधन और जनसंख्या आनुवंशिक गणना शामिल है। यह 3D परमाणु समन्वय हेरफेर के लिए संरचनात्मक विश्लेषण टूल्स, साथ ही जीनोमिक फीचर विज़ुअलाइज़ेशन और जैव-भौगोलिक डेटा मॉडलिंग के लिए यूटिलिटीज भी प्रदान करती है। यह सिस्टम टूल रैपिंग के माध्यम से बाहरी बायोइनफॉरमैटिक्स बाइनरीज के साथ एकीकृत होता है और SQL-समर्थित अनुक्रम भंडारण के माध्यम से निरंतर जैविक रिकॉर्ड भंडारण का समर्थन करता है।
Removes noise and unwanted artifacts from raw biological sequence data to improve analysis quality.
klib is a comprehensive C standard library extension and data structure toolkit. It provides a set of fundamental tools for memory management, data organization, and general-purpose utility functions for standalone C applications. The project features specialized capabilities for bioinformatics sequence analysis, including the parsing of FASTA, FASTQ, and Newick formats and the implementation of Smith-Waterman sequence alignment and Hidden Markov Models. It also includes a mathematical computation library for numerical routines and expression evaluation, as well as a lightweight HTTP and FTP
Provides tools for analyzing biological sequences, including FASTA/FASTQ parsing and Smith-Waterman alignment.
Boltz is a deep learning molecular modeler and biomolecular structure prediction system. It uses neural network architectures to simulate the physical folding and docking of biomolecules, specifically predicting the three-dimensional shapes of protein and ligand complexes. The project functions as a protein-ligand complex predictor and binding affinity predictor, estimating the strength and probability of molecular interactions between ligands and targets. These capabilities are applied to computer aided drug design, including ligand binding affinity prediction and protein-ligand interaction
Transforms raw protein sequences into high-dimensional feature vectors using evolutionary information from related sequences.
evo2 एक जीनोमिक लार्ज लैंग्वेज मॉडल और फाउंडेशन मॉडल है जिसे विभिन्न प्रजातियों में आनुवंशिक जानकारी की भविष्यवाणी, जनरेशन और एनालिसिस के लिए डिज़ाइन किया गया है। यह न्यूक्लियोटाइड सीक्वेंस मॉडलर और DNA सीक्वेंस जनरेटर के रूप में काम करता है, जो जीनोमिक डेटा को प्रोसेस करने के लिए ट्रांसफॉर्मर-आधारित सीक्वेंस मॉडलिंग का उपयोग करता है। यह सिस्टम सिंथेटिक DNA जनरेशन की क्षमताएं प्रदान करता है, जिससे बायोलॉजिकल प्रॉम्प्ट्स या प्रजाति-विशिष्ट टैग्स के आधार पर नए आनुवंशिक अनुक्रम बनाए जा सकते हैं। यह जीनोमिक वेरिएंट्स को स्कोर करने और DNA अनुक्रमों के भीतर जैविक गुणों का विश्लेषण करने के लिए न्यूक्लियोटाइड लाइक्लीहुड प्रेडिक्शन भी करता है। यह मॉडल इंटरमीडिएट लेयर्स से हाई-डायमेंशनल रिप्रेजेंटेशन निकालकर जीनोमिक सीक्वेंस एनालिसिस का समर्थन करता है। ये एम्बेडिंग्स आनुवंशिक डेटा के विशेष वर्गीकरण और डाउनस्ट्रीम एनालिसिस को सक्षम बनाती हैं।
Predicts nucleotide likelihoods across a sequence to score genomic variants or analyze biological properties.
Space Station 14 is a C# multiplayer game and roleplay simulation framework. It is built upon an Entity-Component-System (ECS) game engine that separates logic into systems and data into components to manage complex entity interactions. The project functions as a grid-based physics simulator with a YAML data-driven prototype system for defining game objects. The project features a specialized 2D sprite rendering engine that maps server-side appearance data to client-side shaders. It implements a networking model with client-side prediction and dirty-flagged state synchronization to reduce inp
Allows editing base pairs within a plant's genetic sequence via a hex-style interface.