awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
alchaincyf avatar

alchaincyf/darwin-skill

0
View on GitHub↗
4,343 स्टार्स·476 फोर्क्स·HTML·4 व्यूज़

Darwin Skill

यह प्रोजेक्ट LLM एजेंट कौशल के पुनरावृत्त (iterative) ऑप्टिमाइज़ेशन और सत्यापन के लिए एक फ्रेमवर्क है। यह एक एजेंट क्षमता ऑर्केस्ट्रेटर और प्रॉम्प्ट ऑप्टिमाइज़र के रूप में कार्य करता है, जो भारित रूब्रिक्स और स्वचालित रीराइटिंग के माध्यम से प्रदर्शन को मापने के लिए एक मूल्यांकन फ्रेमवर्क का उपयोग करता है।

यह सिस्टम एक क्लोज्ड-लूप ऑप्टिमाइज़ेशन चक्र के माध्यम से खुद को अलग करता है जो एंकरिंग प्रभावों को रोकने के लिए स्वतंत्र समीक्षक एजेंटों को नियोजित करता है और एक रैचेट-आधारित संस्करण नियंत्रण तंत्र जो स्वचालित रूप से परिवर्तनों को वापस कर देता है यदि वे बेसलाइन स्कोर में सुधार करने में विफल रहते हैं। इसमें वृद्धिशील ट्यूनिंग पठार (plateaus) होने पर स्थानीय इष्टतम (local optima) को दूर करने के लिए खोजपूर्ण संरचनात्मक रीराइटिंग की सुविधा भी है।

प्लेटफ़ॉर्म बहु-आयामी कौशल मूल्यांकन, निष्पादन लिफ्ट को मापने के लिए प्रदर्शन बेंचमार्किंग, और कोड डिफ़्स के मैन्युअल सत्यापन के लिए मानव-इन-द-लूप ओवरसाइट गेट्स सहित व्यापक क्षमताओं को कवर करता है। यह स्टेट-आधारित इतिहास ट्रैकिंग और दृश्य प्रदर्शन कार्ड के निर्माण के माध्यम से ऑब्जर्वेबिलिटी बनाए रखता है।

Features

  • Evaluator-Optimizer Loops - Provides a closed-loop evaluator-optimizer loop to iteratively refine agent skill definitions against structured rubrics.
  • AI Skill Evaluations - Provides a framework for measuring the accuracy and performance of specific AI agent skills against multi-dimensional rubrics.
  • Skill Optimization - Refines agent capabilities by identifying low-scoring dimensions and applying targeted improvements.
  • Automated Prompt Optimization - Iteratively refines model instructions and examples based on multi-dimensional performance metrics.
  • Capability Orchestrators - Manages the lifecycle and versioning of standardized skill definitions across agent environments.
  • Independent Agent Validators - Implements a multi-agent validation pipeline using independent reviewers to ensure unbiased performance gains.
  • LLM Agent Optimization Platforms - Provides a comprehensive platform for developing, evaluating, and monitoring the performance of LLM agent skills.
  • LLM Evaluation Frameworks - Quantifies skill quality using weighted rubrics, static structural analysis, and live execution testing.
  • Baseline Sampling Comparisons - Implements baseline sampling comparisons to quantify the performance lift provided by specific agent skills.
  • Multi-Agent Analysis Pipelines - Utilizes independent reviewer agents in an analysis pipeline to verify performance gains and prevent anchoring effects.
  • Multi-Dimensional Scoring - Employs multi-dimensional scoring with weighted parameters to evaluate skill quality across various structural and functional dimensions.
  • Self-Improving Logic - Automates the optimization of AI instructions through iterative cycles of evaluation and targeted refinement.
  • Score-Based Rollbacks - Employs a ratchet-based version control mechanism that automatically reverts modifications failing to improve baseline scores.
  • Agent Performance Benchmarks - Quantifies the execution lift of specific skills by comparing agent outputs in standardized benchmarking environments.
  • AI Skill Benchmarking - Quantifies the actual execution lift by running identical test prompts with and without specific agent skills.
  • Human-in-the-Loop Oversight - Provides a checkpoint system for human oversight and manual review of performance reports within autonomous workflows.
  • Automated Prompt Engineering - Programmatically refines and optimizes prompt instructions using a structured rubric and automated testing.
  • Human-in-the-loop Controls - Ships mechanisms that pause automated optimization to require manual verification of code diffs and score changes.
  • Human-in-the-Loop Systems - Integrates manual review and validation gates into autonomous optimization cycles to prevent regressions.
  • Structural Prompt Rewriting - Triggers total structural overhauls of skill definitions when incremental tuning plateaus to overcome local optima.
  • Runtime Compatibility Validation - Ensures skill definitions remain agnostic to specific runtimes and avoid platform-specific phrasing.
  • Performance Ratchets - Implements a ratchet-based version control system that reverts modifications if they fail to improve baseline performance scores.
  • Optimization Histories - Maintains a detailed state-based history of commit hashes and score changes to monitor skill evolution.
  • Agent Tooling - Autonomous system for optimizing and evolving agent skill quality.

स्टार हिस्ट्री

alchaincyf/darwin-skill के लिए स्टार हिस्ट्री चार्टalchaincyf/darwin-skill के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

Darwin Skill के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Darwin Skill के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • kiln-ai/kilnkiln-ai का अवतार

    kiln-ai/kiln

    4,910GitHub पर देखें↗

    Kiln is an LLM development workbench and evaluation framework designed for designing, testing, and optimizing prompts and AI agents. It functions as a multi-agent orchestrator and a RAG optimization tool, providing a visual interface for the iterative development of AI systems. The project distinguishes itself through a comprehensive fine-tuning pipeline that supports zero-code model training and reasoning distillation. It enables the creation of hierarchical multi-agent systems where specialized actors coordinate via tool calling, and it implements a Model Context Protocol server to expose t

    Python
    GitHub पर देखें↗4,910
  • coze-dev/coze-loopcoze-dev का अवतार

    coze-dev/coze-loop

    5,540GitHub पर देखें↗

    Coze-loop is an optimization platform and orchestration management suite for large language model agents. It functions as a comprehensive environment for the development, debugging, evaluation, and monitoring of AI agent performance. The project provides a dedicated prompt engineering playground for real-time iteration and validation of model responses. It includes an evaluation framework that runs automated assessments against datasets to generate performance metrics and verify output accuracy. The system covers observability through real-time execution tracing and historical analysis of ag

    Goagentagent-evaluationagent-observability
    GitHub पर देखें↗5,540
  • microsoft/promptwizardmicrosoft का अवतार

    microsoft/PromptWizard

    3,888GitHub पर देखें↗

    PromptWizard is an automated prompt engineering framework designed to evolve natural language instructions for generative tasks. It functions as an in-context learning optimizer and synthetic data generator, using mutation rounds and performance metrics to iteratively refine large language model instructions. The system employs a self-reflective optimization loop that uses model-generated critiques to rewrite prompts. It distinguishes itself through the use of reasoning chain integration and persona-based prompting to steer the tone and professional quality of model responses. The framework

    Python
    GitHub पर देखें↗3,888
  • keirp/automatic_prompt_engineerkeirp का अवतार

    keirp/automatic_prompt_engineer

    1,360GitHub पर देखें↗

    Automatic Prompt Engineer is a framework designed to automate the generation, refinement, and performance measurement of language model instructions. It functions as a systematic tool for optimizing prompt phrasing by iteratively testing candidate instructions against specific input and output datasets to maximize task accuracy. The system distinguishes itself through an evaluation-driven approach that uses automated feedback loops to score prompt variations. By employing template-based input structuring, it ensures consistent testing environments where candidate instructions are measured aga

    Python
    GitHub पर देखें↗1,360
Darwin Skill के सभी 30 विकल्प देखें→

अक्सर पूछे जाने वाले प्रश्न

alchaincyf/darwin-skill क्या करता है?

यह प्रोजेक्ट LLM एजेंट कौशल के पुनरावृत्त (iterative) ऑप्टिमाइज़ेशन और सत्यापन के लिए एक फ्रेमवर्क है। यह एक एजेंट क्षमता ऑर्केस्ट्रेटर और प्रॉम्प्ट ऑप्टिमाइज़र के रूप में कार्य करता है, जो भारित रूब्रिक्स और स्वचालित रीराइटिंग के माध्यम से प्रदर्शन को मापने के लिए एक मूल्यांकन फ्रेमवर्क का उपयोग करता है।

alchaincyf/darwin-skill की मुख्य विशेषताएं क्या हैं?

alchaincyf/darwin-skill की मुख्य विशेषताएं हैं: Evaluator-Optimizer Loops, AI Skill Evaluations, Skill Optimization, Automated Prompt Optimization, Capability Orchestrators, Independent Agent Validators, LLM Agent Optimization Platforms, LLM Evaluation Frameworks।

alchaincyf/darwin-skill के कुछ ओपन-सोर्स विकल्प क्या हैं?

alchaincyf/darwin-skill के ओपन-सोर्स विकल्पों में शामिल हैं: kiln-ai/kiln — Kiln is an LLM development workbench and evaluation framework designed for designing, testing, and optimizing prompts… coze-dev/coze-loop — Coze-loop is an optimization platform and orchestration management suite for large language model agents. It functions… microsoft/promptwizard — PromptWizard is an automated prompt engineering framework designed to evolve natural language instructions for… keirp/automatic_prompt_engineer — Automatic Prompt Engineer is a framework designed to automate the generation, refinement, and performance measurement… mshumer/gpt-prompt-engineer — This project is an automated prompt engineering and optimization tool designed to iteratively create, test, and refine… stanfordnlp/dspy — DSPy is a declarative programming framework designed for building complex language model applications. It treats model…