awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
dariusk avatar

dariusk/corpora

0
View on GitHub↗
5,099 Stars·1,296 Forks·JavaScript·5 Aufrufe

Corpora

Corpora is a library of curated, domain-specific text corpora and structured datasets. It provides collections of categorized nouns, adjectives, and verbs in JSON format, intended for use as static training or testing data for conversational agents and automated systems.

The project organizes linguistic data across diverse fields, including science, art, and geography. These datasets are distributed as standalone files using a language-neutral schema and a static JSON data model, allowing for direct import into applications without an API layer.

The library supports chatbot content generation and rapid prototype development by providing structured data integration. It utilizes part-of-speech categorization and domain-specific data partitioning to facilitate the creation of dialogue and the testing of machine learning models.

Features

  • Training Corpora - Provides focused text corpora used as static training or testing data for conversational agents.
  • AI Chatbots - Enables the creation of diverse dialogue and responses for bots using curated linguistic lists.
  • Domain-Specific Pre-training Corpora - Offers curated datasets covering diverse fields such as science, art, and geography for language-based applications.
  • Static Training Sets - Provides focused collections of specialized terms in science and art to test or train small machine learning models.
  • Grammatical Category Filters - Groups words by grammatical categories such as nouns, verbs, and adjectives for structured text generation.
  • Curated Datasets - Provides access to small, focused collections of categorized words in JSON format for prototyping.
  • JSON Dataset Collections - Provides a set of small, structured JSON files containing categorized nouns, adjectives, and verbs.
  • Static JSON Stores - Stores structured linguistic categories in read-only JSON files for immediate retrieval without a database.
  • Flat-File Storage - Distributes datasets as standalone text files to allow direct import without an API layer.
  • Rapid Prototyping Environments - Facilitates building early application versions quickly using structured JSON datasets instead of full databases.
  • Knowledge Domain Partitioning - Organizes linguistic datasets into topical categories to ensure broad knowledge coverage for conversational agents.
  • Language-Neutral Schemas - Uses a consistent data format across collections to ensure compatibility with various bot development frameworks.

Star-Verlauf

Star-Verlauf für dariusk/corporaStar-Verlauf für dariusk/corpora

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Corpora

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Corpora.
  • samayo/country-jsonAvatar von samayo

    samayo/country-json

    1,148Auf GitHub ansehen↗

    Country-json is a repository of standardized international metadata providing machine-readable datasets that encompass comprehensive geographic, demographic, and cultural country attributes. It functions as a collection of structured information for every country in the world, formatted as plain text files to ensure universal compatibility across programming environments. The project utilizes a decoupled distribution model that relies on static file-system retrieval rather than database connections or network requests. By organizing global information into standardized key-value pairs, it all

    JavaScriptcapital-citycitycountries
    Auf GitHub ansehen↗1,148
  • kkndmetianya/kkndme_tianyaAvatar von kkndmetianya

    kkndmetianya/kkndme_tianya

    19,404Auf GitHub ansehen↗

    This project is a static site forum archive and housing market dataset. It serves as a read-only repository of preserved community discussions focused on residential property pricing and real estate trends, stored as a collection of Markdown files to ensure long-term data preservation and portability. The system functions as a web content preservation tool that converts historical forum threads into pre-rendered HTML files for serverless access. This approach organizes forum content archiving and housing market analysis into a permanent digital format to prevent data loss. The archive is bui

    Auf GitHub ansehen↗19,404
  • rfordatascience/tidytuesdayAvatar von rfordatascience

    rfordatascience/tidytuesday

    8,211Auf GitHub ansehen↗

    This repository provides a curated collection of weekly datasets designed for data visualization practice, data science education, and statistical analysis. It serves as a central source for cleaned and structured real-world data, allowing practitioners to focus on analysis and visualization without the need to scrape or clean raw files. The project facilitates a community learning workflow where users can explore a wide variety of topics, ranging from global health spending and energy datasets to maritime logs and baby name popularity. Participants are encouraged to share their resulting vis

    HTML
    Auf GitHub ansehen↗8,211
  • opencx-labs/openchatAvatar von opencx-labs

    opencx-labs/OpenChat

    5,264Auf GitHub ansehen↗

    OpenChat is a conversational AI agent builder and customer service automation platform that uses large language models to power customer support chatbots across multiple channels. It provides tools for defining AI agent behavior, training on custom knowledge, managing actions, and controlling autopilot responses per channel. The platform enables deploying AI agents on web, phone, email, SMS, and WhatsApp, with a unified inbox for managing conversations across all channels. It includes CRM synchronization, automated workflows, contact segmentation, and analytics for tracking customer satisfact

    JavaScript
    Auf GitHub ansehen↗5,264
Alle 30 Alternativen zu Corpora anzeigen→

Häufig gestellte Fragen

Was macht dariusk/corpora?

Corpora is a library of curated, domain-specific text corpora and structured datasets. It provides collections of categorized nouns, adjectives, and verbs in JSON format, intended for use as static training or testing data for conversational agents and automated systems.

Was sind die Hauptfunktionen von dariusk/corpora?

Die Hauptfunktionen von dariusk/corpora sind: Training Corpora, AI Chatbots, Domain-Specific Pre-training Corpora, Static Training Sets, Grammatical Category Filters, Curated Datasets, JSON Dataset Collections, Static JSON Stores.

Welche Open-Source-Alternativen gibt es zu dariusk/corpora?

Open-Source-Alternativen zu dariusk/corpora sind unter anderem: samayo/country-json — Country-json is a repository of standardized international metadata providing machine-readable datasets that encompass… kkndmetianya/kkndme_tianya — This project is a static site forum archive and housing market dataset. It serves as a read-only repository of… rfordatascience/tidytuesday — This repository provides a curated collection of weekly datasets designed for data visualization practice, data… vercel-labs/ai-chatbot — This is a full-featured chatbot framework and Next.js web application designed for integrating various large language… opencx-labs/openchat — OpenChat is a conversational AI agent builder and customer service automation platform that uses large language models… memochou1993/gpt-ai-assistant — This project is a serverless application that integrates OpenAI models with the LINE messaging platform. It functions…