7 个仓库
Specialized databases and tools for scalable genomic analysis.
Explore 7 awesome GitHub repositories matching part of an awesome list · Genomic Data Analysis. Refine with filters or upvote what's useful.
Accelerates standard genomics workflows using GPU-optimized versions of open-source tools.
Biopython 是一个 Python 生物信息学库,提供用于解析、操作和分析生物序列、分子结构和系统发育树的工具。它作为基因组和蛋白质组数据的生物序列解析器,支持多种行业标准文件格式,并充当从 NCBI Entrez 仓库查询生物数据和引用的接口。 该项目以其用于蛋白质结构分析和系统发育树构建的专业工具包而著称。它包括用于处理 PDB 和 mmCIF 文件以计算分子几何结构的蛋白质结构分析器,以及用于分析物种间进化关系的系统发育树工具包。 该库涵盖了广泛的生物信息学能力,包括用于转录和翻译的基因组序列分析、序列比对管理以及群体遗传学计算。它还提供用于 3D 原子坐标操作的结构分析工具,以及用于基因组特征可视化和生物地理数据建模的实用程序。 该系统通过工具封装与外部生物信息学二进制文件集成,并支持通过 SQL 后端进行持久化生物记录存储。
Parses and manipulates DNA and RNA sequences to identify features and generate reverse complements.
evo2 是一个基因组大语言模型和基础模型,旨在预测、生成和分析不同物种的遗传信息。它作为一个核苷酸序列建模器和 DNA 序列生成器,使用基于 Transformer 的序列建模来处理基因组数据。 该系统提供了合成 DNA 生成功能,可根据生物学提示或物种特定标签创建新的遗传序列。它还执行核苷酸可能性预测,以对基因组变异进行评分并分析 DNA 序列中的生物学特性。 该模型通过从中间层提取高维表示来支持基因组序列分析。这些嵌入使得能够对遗传数据进行专门的分类和下游分析。
Extracts high-dimensional embeddings from genetic data to perform specialized biological classification and sequence analysis.
DeepVariant is a deep learning genotyping tool and DNA sequence analysis pipeline used to identify single nucleotide polymorphisms and indels from next-generation sequencing data. It functions as a convolutional neural network genetic variant caller that treats genomic read alignments as multi-channel image tensors to determine genotypes. The system supports specialized analysis workflows including long-read variant calling for circular consensus sequencing and trio-based variant calling to identify inherited or de novo mutations. It enables model optimization for new species or genome contex
Predicts inherited or de novo mutations by calling variants across related family samples.
Scanpy is a Python library for the preprocessing, visualization, and analysis of large-scale single-cell gene expression datasets. It serves as a toolkit for single-cell RNA sequencing analysis, providing a framework to process and analyze genomic data from individual cells to identify biological markers and cell types. The library includes a scalable data processing pipeline for cleaning and preparing genomic data, a clustering framework for grouping cells with similar expression profiles, and a system for modeling transitions between cell states to reconstruct biological development and dif
Provides memory-efficient pipelines for cleaning and preparing large-scale single-cell datasets.
Cloud-native genomic dataframes and batch computing
Platform for scalable genomic data analysis.
Scalable gVCF merging and joint variant calling for population sequencing projects
Tool for scalable gVCF merging and joint variant calling.