awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to asyml/texar

Open-source alternatives to Texar

30 open-source projects similar to asyml/texar, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Texar alternative.

  • huggingface/transformershuggingface 的头像

    huggingface/transformers

    161,630在 GitHub 上查看↗

    Transformers is a comprehensive library for machine learning that provides a unified interface for training, fine-tuning, and deploying transformer-based models. It supports a wide range of tasks, including text classification, language modeling, question answering, and sequence-to-sequence translation, while offering specialized architectures for both text and vision processing. The framework includes tools for managing the entire model lifecycle, from data preprocessing and tokenization to distributed training and inference. The library features extensive support for model optimization and

    Pythonaudiodeep-learningdeepseek
    在 GitHub 上查看↗161,630
  • codertimo/bert-pytorchcodertimo 的头像

    codertimo/BERT-pytorch

    6,518在 GitHub 上查看↗
    Pythonbertlanguage-modelnlp
    在 GitHub 上查看↗6,518
  • facebookresearch/xlmfacebookresearch 的头像

    facebookresearch/XLM

    2,930在 GitHub 上查看↗

    PyTorch original implementation of Cross-lingual Language Model Pretraining.

    Python
    在 GitHub 上查看↗2,930
  • wb14123/couplet-datasetwb14123 的头像

    wb14123/couplet-dataset

    745在 GitHub 上查看↗

    Dataset for couplets. 70万条对联数据库。

    Pythondataset
    在 GitHub 上查看↗745
  • openbmb/bmlistOpenBMB 的头像

    OpenBMB/BMList

    345在 GitHub 上查看↗

    A List of Big Models

    Pythonaiapicode
    在 GitHub 上查看↗345

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • minimaxir/textgenrnnM

    minimaxir/textgenrnn

    0在 GitHub 上查看↗

    Easily train your own text-generating neural network of any size and complexity on any text dataset with a few lines of code, or quickly train on a text using a pretrained model.

    在 GitHub 上查看↗0
  • yannvgn/laserembeddingsyannvgn 的头像

    yannvgn/laserembeddings

    225在 GitHub 上查看↗

    LASER multilingual sentence embeddings as a pip package

    Python
    在 GitHub 上查看↗225
  • nlpscott/bert-chinese-classification-taskNLPScott 的头像

    NLPScott/bert-Chinese-classification-task

    737在 GitHub 上查看↗

    bert中文分类实践

    Python
    在 GitHub 上查看↗737
  • socialbird-ailab/bert-classification-tutorialSocialbird-AILab 的头像

    Socialbird-AILab/BERT-Classification-Tutorial

    535在 GitHub 上查看↗

    标注数据,可以说是AI模型训练里最艰巨的一项工作了。自然语言处理的数据标注更是需要投入大量人力。相对计算机视觉的图像标注,文本的标注通常没有准确的标准答案,对句子理解也是因人而异,让这项工作更是难上加难。 但是!谷歌最近发布的BERT大大的解决了这个问题!根据我们的实验,BERT在文本多分类的任务中,能在极小的数据下,带来显著的分类准确率提升。并且,实验主要对比的是仅仅5个月前发布的State of the art 语言模型迁移学习模型 - ULMFiT (https://arxiv.org/abs/1801.06146), 结果有着明显的提升。

    Python
    在 GitHub 上查看↗535
  • turtlesoupy/this-word-does-not-existturtlesoupy 的头像

    turtlesoupy/this-word-does-not-exist

    1,020在 GitHub 上查看↗

    This Word Does Not Exist

    Python
    在 GitHub 上查看↗1,020
  • terrifyzhao/bert-utilsterrifyzhao 的头像

    terrifyzhao/bert-utils

    1,670在 GitHub 上查看↗

    一行代码使用BERT生成句向量,BERT做文本分类、文本相似度计算

    Python
    在 GitHub 上查看↗1,670
  • huggingface/coursehuggingface 的头像

    huggingface/course

    3,715在 GitHub 上查看↗

    This project is an educational course and learning curriculum for implementing and fine-tuning transformer models using the Hugging Face ecosystem. It serves as a structured guide and technical walkthrough for processing multimodal data, adapting pre-trained neural networks, and deploying models. The material includes a guide for managing, versioning, and distributing model weights and datasets through a centralized asset hub. It also provides a practical tutorial on adapting models to specific datasets using parameter-efficient methods and an implementation guide for solving natural language

    MDXdeep-learninghacktoberfestnlp
    在 GitHub 上查看↗3,715
  • openai/gpt-2openai 的头像

    openai/gpt-2

    24,967在 GitHub 上查看↗

    This project is a transformer-based language model and autoregressive text generator designed to predict the next token in a sequence to produce human-like prose and synthetic text. It functions as a large language model that utilizes a transformer architecture to learn linguistic patterns from large datasets for unsupervised multitask learning. The repository provides a distribution of pre-trained weights, enabling natural language processing tasks without requiring additional training. This allows the model to perform zero-shot task generalization by applying learned patterns to new tasks.

    Python
    在 GitHub 上查看↗24,967
  • facebookresearch/flow_matchingfacebookresearch 的头像

    facebookresearch/flow_matching

    4,562在 GitHub 上查看↗

    This project is a PyTorch-based generative model framework designed to transform noise into complex data distributions by learning vector fields and probability paths. It serves as a multimodal generative toolkit for producing synthetic text and images through learned probability flows. The library distinguishes itself by supporting continuous, discrete, and Riemannian manifold integrations. This allows the framework to handle a variety of data types, including categorical data via discrete-state flow matching and non-Euclidean spaces through Riemannian manifold integration. The toolkit cove

    Python
    在 GitHub 上查看↗4,562
  • morvanzhou/tutorialsMorvanZhou 的头像

    MorvanZhou/tutorials

    12,952在 GitHub 上查看↗

    This repository is a comprehensive collection of instructional guides and practical examples for Python development, focusing on machine learning, data science, and web scraping. It provides implementations for neural networks, reinforcement learning algorithms, and deep learning architectures using PyTorch, alongside detailed manuals for scientific computing and data visualization. The project distinguishes itself by offering specialized tutorials on concurrent programming to optimize CPU performance and guides for setting up Linux development environments. It covers the implementation of ad

    Pythonmachine-learningmultiprocessingneural-network
    在 GitHub 上查看↗12,952
  • alexsergivan/transliteratorA

    alexsergivan/transliterator

    0在 GitHub 上查看↗
    在 GitHub 上查看↗0
  • alexrozanski/llamachatalexrozanski 的头像

    alexrozanski/LlamaChat

    1,510在 GitHub 上查看↗

    Chat with your favourite LLaMA models in a native macOS app

    Swiftaillamallamacpp
    在 GitHub 上查看↗1,510
  • abosamoor/polyglotaboSamoor 的头像

    aboSamoor/polyglot

    2,367在 GitHub 上查看↗

    Multilingual text (NLP) processing toolkit

    Python
    在 GitHub 上查看↗2,367
  • akanimax/natural-language-summary-generation-from-structured-dataakanimax 的头像

    akanimax/natural-language-summary-generation-from-structured-data

    186在 GitHub 上查看↗

    Implementation (Personal) of the paper titled "Order-Planning Neural Text Generation From Structured Data". The dataset for this project can be found at -> WikiBio

    Python
    在 GitHub 上查看↗186
  • aigc-audio/audiogptAIGC-Audio 的头像

    AIGC-Audio/AudioGPT

    10,174在 GitHub 上查看↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Pythonaudiogptmusic
    在 GitHub 上查看↗10,174
  • abitdodgy/gibranabitdodgy 的头像

    abitdodgy/gibran

    65在 GitHub 上查看↗

    Gibran is an Elixir natural language processor, and a port of WordsCounted.

    Elixir
    在 GitHub 上查看↗65
  • 7compass/sentimental7compass 的头像

    7compass/sentimental

    465在 GitHub 上查看↗

    Simple sentiment analysis with Ruby

    Ruby
    在 GitHub 上查看↗465
  • alvations/annotate-questionnairealvations 的头像

    alvations/annotate-questionnaire

    59在 GitHub 上查看↗

    Summary of Responses to Questionnaire on Annotation Platform https://forms.gle/iZk8kehkjAWmB8xe9

    在 GitHub 上查看↗59
  • anujvyas/natural-language-processing-projectsanujvyas 的头像

    anujvyas/Natural-Language-Processing-Projects

    254在 GitHub 上查看↗

    This repository consists of all my NLP Projects

    Jupyter Notebook
    在 GitHub 上查看↗254
  • arc53/docsgptarc53 的头像

    arc53/DocsGPT

    17,939在 GitHub 上查看↗

    DocsGPT is a retrieval-augmented generation platform and private knowledge base used to build AI agents that perform grounded search and analysis. It functions as a multi-model AI orchestrator and enterprise agent builder, allowing for the integration of various local and cloud language models to customize reasoning and text generation. The project provides a visual environment for developing automated assistants using conditional logic and third-party API connectivity. It enables the creation of private AI agents capable of performing enterprise search and detailed document analysis using pr

    Pythonagent-builderagentsai
    在 GitHub 上查看↗17,939
  • argilla-io/argillaargilla-io 的头像

    argilla-io/argilla

    5,015在 GitHub 上查看↗

    Argilla is a collaborative AI feedback tool and data curation management system. It serves as a human-in-the-loop dataset platform designed to coordinate workforce annotators and domain experts in labeling, rating, and refining data samples for machine learning projects. The platform focuses on large language model dataset curation and reinforcement learning from human feedback workflows. It provides a shared workspace for integrating human expertise into AI development to validate model outputs and correct data errors. The system manages the end-to-end machine learning data pipeline, includ

    Python
    在 GitHub 上查看↗5,015
  • arongdari/python-topic-modelarongdari 的头像

    arongdari/python-topic-model

    374在 GitHub 上查看↗

    Implementation of various topic models

    Jupyter Notebook
    在 GitHub 上查看↗374
  • arongdari/topic-model-lecture-notearongdari 的头像

    arongdari/topic-model-lecture-note

    22在 GitHub 上查看↗

    lecture notes for probabilistic topic models using ipython notebook

    在 GitHub 上查看↗22
  • artidoro/qloraartidoro 的头像

    artidoro/qlora

    10,929在 GitHub 上查看↗

    This project is a quantized fine-tuning framework for large language models. It implements a low-rank adaptation library and a four-bit quantizer to reduce the GPU memory requirements needed to train large models. The framework utilizes four-bit quantization and low-rank adapters to enable model training on consumer-grade hardware. It further reduces the memory footprint through double quantization and a paged optimizer that offloads states to system RAM. The system supports distributed training across multiple GPUs to handle larger parameter scales and includes utilities for custom dataset

    Jupyter Notebook
    在 GitHub 上查看↗10,929
  • allenai/scispacyallenai 的头像

    allenai/SciSpaCy

    1,968在 GitHub 上查看↗

    This repository contains custom pipes and models related to using spaCy for scientific documents.

    Python
    在 GitHub 上查看↗1,968