awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
dariusk avatar

dariusk/corpora

0
View on GitHub↗
5,099 stars·1,296 forks·JavaScript·12 views

Corpora

Corpora is a library of curated, domain-specific text corpora and structured datasets. It provides collections of categorized nouns, adjectives, and verbs in JSON format, intended for use as static training or testing data for conversational agents and automated systems.

The project organizes linguistic data across diverse fields, including science, art, and geography. These datasets are distributed as standalone files using a language-neutral schema and a static JSON data model, allowing for direct import into applications without an API layer.

The library supports chatbot content generation and rapid prototype development by providing structured data integration. It utilizes part-of-speech categorization and domain-specific data partitioning to facilitate the creation of dialogue and the testing of machine learning models.

Features

  • Training Corpora - Provides focused text corpora used as static training or testing data for conversational agents.
  • AI Chatbots - Enables the creation of diverse dialogue and responses for bots using curated linguistic lists.
  • Domain-Specific Pre-training Corpora - Offers curated datasets covering diverse fields such as science, art, and geography for language-based applications.
  • Static Training Sets - Provides focused collections of specialized terms in science and art to test or train small machine learning models.
  • Grammatical Category Filters - Groups words by grammatical categories such as nouns, verbs, and adjectives for structured text generation.
  • Curated Datasets - Provides access to small, focused collections of categorized words in JSON format for prototyping.
  • JSON Dataset Collections - Provides a set of small, structured JSON files containing categorized nouns, adjectives, and verbs.
  • Static JSON Stores - Stores structured linguistic categories in read-only JSON files for immediate retrieval without a database.
  • Flat-File Storage - Distributes datasets as standalone text files to allow direct import without an API layer.
  • Rapid Prototyping Environments - Facilitates building early application versions quickly using structured JSON datasets instead of full databases.
  • Knowledge Domain Partitioning - Organizes linguistic datasets into topical categories to ensure broad knowledge coverage for conversational agents.
  • Language-Neutral Schemas - Uses a consistent data format across collections to ensure compatibility with various bot development frameworks.

Star history

Star history chart for dariusk/corporaStar history chart for dariusk/corpora

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Corpora

These projects share indexed features with Corpora. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • samayo/country-jsonsamayo avatar

    samayo/country-json

    1,148View on GitHub↗

    Country-json is a repository of standardized international metadata providing machine-readable datasets that encompass comprehensive geographic, demographic, and cultural country attributes. It functions as a collection of structured information for every country in the world, formatted as plain text files to ensure universal compatibility across programming environments. The project utilizes a decoupled distribution model that relies on static file-system retrieval rather than database connections or network requests. By organizing global information into standardized key-value pairs, it all

    JavaScriptcapital-citycitycountries
    View on GitHub↗1,148
  • kkndmetianya/kkndme_tianyakkndmetianya avatar

    kkndmetianya/kkndme_tianya

    19,404View on GitHub↗

    This project is a static site forum archive and housing market dataset. It serves as a read-only repository of preserved community discussions focused on residential property pricing and real estate trends, stored as a collection of Markdown files to ensure long-term data preservation and portability. The system functions as a web content preservation tool that converts historical forum threads into pre-rendered HTML files for serverless access. This approach organizes forum content archiving and housing market analysis into a permanent digital format to prevent data loss. The archive is bui

    View on GitHub↗19,404
  • rfordatascience/tidytuesdayrfordatascience avatar

    rfordatascience/tidytuesday

    8,211View on GitHub↗

    This repository provides a curated collection of weekly datasets designed for data visualization practice, data science education, and statistical analysis. It serves as a central source for cleaned and structured real-world data, allowing practitioners to focus on analysis and visualization without the need to scrape or clean raw files. The project facilitates a community learning workflow where users can explore a wide variety of topics, ranging from global health spending and energy datasets to maritime logs and baby name popularity. Participants are encouraged to share their resulting vis

    HTML
    View on GitHub↗8,211
  • opencx-labs/openchatopencx-labs avatar

    opencx-labs/OpenChat

    5,264View on GitHub↗

    OpenChat is a conversational AI agent builder and customer service automation platform that uses large language models to power customer support chatbots across multiple channels. It provides tools for defining AI agent behavior, training on custom knowledge, managing actions, and controlling autopilot responses per channel. The platform enables deploying AI agents on web, phone, email, SMS, and WhatsApp, with a unified inbox for managing conversations across all channels. It includes CRM synchronization, automated workflows, contact segmentation, and analytics for tracking customer satisfact

    JavaScript
    View on GitHub↗5,264
Compare all 30 related projects→

Frequently asked questions

What does dariusk/corpora do?

Corpora is a library of curated, domain-specific text corpora and structured datasets. It provides collections of categorized nouns, adjectives, and verbs in JSON format, intended for use as static training or testing data for conversational agents and automated systems.

What are the main features of dariusk/corpora?

The main features of dariusk/corpora are: Training Corpora, AI Chatbots, Domain-Specific Pre-training Corpora, Static Training Sets, Grammatical Category Filters, Curated Datasets, JSON Dataset Collections, Static JSON Stores.

Which projects share features with dariusk/corpora?

Projects with overlapping indexed features include: samayo/country-json — Country-json is a repository of standardized international metadata providing machine-readable datasets that encompass… kkndmetianya/kkndme_tianya — This project is a static site forum archive and housing market dataset. It serves as a read-only repository of… rfordatascience/tidytuesday — This repository provides a curated collection of weekly datasets designed for data visualization practice, data… vercel-labs/ai-chatbot — This is a full-featured chatbot framework and Next.js web application designed for integrating various large language… opencx-labs/openchat — OpenChat is a conversational AI agent builder and customer service automation platform that uses large language models… memochou1993/gpt-ai-assistant — This project is a serverless application that integrates OpenAI models with the LINE messaging platform. It functions…