awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

11 Repos

Awesome GitHub RepositoriesSafety and Alignment Frameworks

Tools and validation layers for ensuring generative model outputs adhere to safety guidelines and organizational standards.

Distinguishing note: Focuses on the governance and filtering layer of AI models rather than the training or inference infrastructure.

Explore 11 awesome GitHub repositories matching artificial intelligence & ml · Safety and Alignment Frameworks. Refine with filters or upvote what's useful.

Awesome Safety and Alignment Frameworks GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • rohitg00/ai-engineering-from-scratchAvatar von rohitg00

    rohitg00/ai-engineering-from-scratch

    33,575Auf GitHub ansehen↗

    This project is a structured AI engineering curriculum and educational program designed to teach the construction of machine learning models, neural networks, and autonomous agents from the ground up. It serves as a comprehensive machine learning course covering mathematical foundations, deep learning architectures, and reinforcement learning through practical implementation. The project provides a technical framework for building autonomous loops and memory systems via an agent framework, as well as guides for implementing multimodal AI systems that integrate vision, audio, and text processi

    Implements methodologies for red-teaming, constitutional AI, and reward model training to ensure model safety.

    Pythonagentsaiai-agents
    Auf GitHub ansehen↗33,575
  • meta-llama/llama3Avatar von meta-llama

    meta-llama/llama3

    29,254Auf GitHub ansehen↗

    Llama 3 is a collection of pretrained, autoregressive transformer-based models designed for natural language generation, reasoning, and complex instruction following. It functions as a generative AI framework that provides the infrastructure for managing model weights, executing neural network inference, and handling computational workloads across diverse knowledge domains. The project distinguishes itself through an integrated AI safety toolkit that employs secondary classification filtering to inspect inputs and outputs, ensuring adherence to usage compliance and safety standards. It suppor

    Implements secondary classification filtering to inspect inputs and outputs, ensuring adherence to safety and usage compliance standards.

    Python
    Auf GitHub ansehen↗29,254
  • elder-plinius/l1b3rt4sAvatar von elder-plinius

    elder-plinius/L1B3RT4S

    20,033Auf GitHub ansehen↗

    L1B3RT4S is an adversarial machine learning toolkit designed for red teaming and evaluating the robustness of large language models. It provides a research framework for investigating how safety alignment mechanisms and content moderation systems respond to sophisticated input strategies. The project focuses on identifying vulnerabilities in model guardrails by employing techniques such as adversarial narrative framing, dynamic context injection, and latent space steering. It utilizes multi-agent prompt decomposition and recursive text transformation to analyze how structural changes to input

    Investigates the robustness of alignment mechanisms by testing how various prompt engineering techniques influence model output and safety filter behavior.

    1337adversarial-attacksai
    Auf GitHub ansehen↗20,033
  • nvidia-nemo/nemoAvatar von NVIDIA-NeMo

    NVIDIA-NeMo/NeMo

    17,389Auf GitHub ansehen↗

    NeMo is a comprehensive framework designed for the development, training, and deployment of large-scale conversational and generative artificial intelligence models. It provides an integrated platform for building multimodal systems, encompassing speech processing, language modeling, and reinforcement learning alignment. The framework is built to handle the entire lifecycle of AI development, from data curation and model pretraining to production-ready service deployment. The platform distinguishes itself through advanced distributed training capabilities, including tensor and pipeline parall

    Implements programmable guardrails to enforce safety policies and topical boundaries on model outputs.

    Pythonasrdeeplearninggenerative-ai
    Auf GitHub ansehen↗17,389
  • googlecloudplatform/generative-aiAvatar von GoogleCloudPlatform

    GoogleCloudPlatform/generative-ai

    12,700Auf GitHub ansehen↗

    This project is a development platform for managing the lifecycle of generative artificial intelligence models. It provides a unified environment for accessing, fine-tuning, and deploying large language models, serving as an orchestrator that handles the integration of diverse models into custom applications. The platform distinguishes itself by offering a managed infrastructure for hosting and scaling models, which removes the requirement for manual server maintenance or configuration. It includes integrated tools for supervised fine-tuning and vector embedding optimization, allowing for the

    Apply content filtering and responsible guidelines to monitor and restrict model outputs, ensuring that all generated responses adhere to established safety policies and ethical usage standards.

    Jupyter Notebookagentsgcpgemini
    Auf GitHub ansehen↗12,700
  • wandb/wandbAvatar von wandb

    wandb/wandb

    10,844Auf GitHub ansehen↗

    Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow orchestration. It provides a comprehensive suite of tools for logging, visualizing, and versioning training metrics, model artifacts, and hyperparameter sweeps to ensure reproducibility across development cycles. The platform also functions as an observability tool for large language model applications, enabling the tracing of execution steps, token usage, and reasoning processes. The project distinguishes itself through its event-driven automation capabilities, which allow users

    Evaluates model alignment, bias, and safety to ensure responsible behavior.

    Pythonaicollaborationdata-science
    Auf GitHub ansehen↗10,844
  • ymcui/chinese-llama-alpaca-2Avatar von ymcui

    ymcui/Chinese-LLaMA-Alpaca-2

    7,136Auf GitHub ansehen↗

    This project provides a Chinese large language model based on the LLaMA architecture. It is an instruction-tuned model optimized for natural language processing and multi-turn conversations in Chinese. The system includes a framework for parameter-efficient fine-tuning using low-rank adaptation and quantization to reduce memory requirements. It also implements retrieval augmented generation for local document question answering and supports long-context processing for sequences up to 64K tokens. The project covers a broad set of capabilities including supervised instruction tuning, reinforce

    Implements safety and alignment frameworks to ensure model outputs adhere to ethical guidelines.

    Python64kalpacaalpaca-2
    Auf GitHub ansehen↗7,136
  • google/gemma_pytorchAvatar von google

    google/gemma_pytorch

    5,697Auf GitHub ansehen↗

    The official PyTorch implementation of Google's Gemma models

    Implements a safety classifier that evaluates text against policy-defined categories to detect violations.

    Pythongemmagooglepytorch
    Auf GitHub ansehen↗5,697
  • dontriskit/awesome-ai-system-promptsAvatar von dontriskit

    dontriskit/awesome-ai-system-prompts

    5,206Auf GitHub ansehen↗

    This project is a comprehensive library of structured system prompts and configuration templates designed to define the behavior, persona, and operational boundaries of autonomous artificial intelligence agents. It serves as a framework for prompt engineering, providing modular instructions that help models parse complex tasks, maintain consistent interaction tones, and adhere to specific domain constraints. The repository distinguishes itself by offering specialized configurations for agent safety and security, including protocols to prevent prompt injection and unauthorized data access. It

    Provides frameworks for defining safety boundaries and standardized refusal protocols for AI interactions.

    TypeScript
    Auf GitHub ansehen↗5,206
  • llm-attacks/llm-attacksAvatar von llm-attacks

    llm-attacks/llm-attacks

    4,509Auf GitHub ansehen↗

    This repository provides tools and methodologies for studying adversarial attacks on large language models. It focuses on understanding how carefully crafted inputs can manipulate or bypass the safety mechanisms of LLMs, enabling researchers to probe model vulnerabilities and improve their robustness. The project covers techniques for generating adversarial prompts, evaluating model responses under attack conditions, and analyzing the effectiveness of different attack strategies.

    Provides a framework for identifying alignment failures to ensure models do not generate harmful content.

    Python
    Auf GitHub ansehen↗4,509
  • zjunlp/easyeditAvatar von zjunlp

    zjunlp/EasyEdit

    2,718Auf GitHub ansehen↗

    EasyEdit is a framework and toolkit designed for updating, inserting, or erasing specific factual information within large language models without requiring full retraining. It functions as a parameter modifier and knowledge editing system capable of performing targeted weight updates across diverse model architectures. The project distinguishes itself by supporting both text-based and multimodal model editing, allowing for knowledge updates across image and text modalities. It provides utilities for model steering to adjust personality and reasoning patterns in real time via activation inter

    Locates and edits specific model regions to remove toxic behaviors and neutralize harmful outputs.

    Jupyter Notebookartificial-intelligencebaichuanchatgpt
    Auf GitHub ansehen↗2,718
  1. Home
  2. Artificial Intelligence & ML
  3. Safety and Alignment Frameworks

Unter-Tags erkunden

  • Content Safety ClassifiersModules that evaluate text against predefined safety policies and flag policy-violating content. **Distinct from Safety and Alignment Frameworks:** Distinct from Safety and Alignment Frameworks: focuses on the classification step itself rather than the broader governance and filtering layer.