awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

11 repository-uri

Awesome GitHub RepositoriesSafety and Alignment Frameworks

Tools and validation layers for ensuring generative model outputs adhere to safety guidelines and organizational standards.

Distinguishing note: Focuses on the governance and filtering layer of AI models rather than the training or inference infrastructure.

Explore 11 awesome GitHub repositories matching artificial intelligence & ml · Safety and Alignment Frameworks. Refine with filters or upvote what's useful.

Awesome Safety and Alignment Frameworks GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • rohitg00/ai-engineering-from-scratchAvatar rohitg00

    rohitg00/ai-engineering-from-scratch

    33,575Vezi pe GitHub↗

    This project is a structured AI engineering curriculum and educational program designed to teach the construction of machine learning models, neural networks, and autonomous agents from the ground up. It serves as a comprehensive machine learning course covering mathematical foundations, deep learning architectures, and reinforcement learning through practical implementation. The project provides a technical framework for building autonomous loops and memory systems via an agent framework, as well as guides for implementing multimodal AI systems that integrate vision, audio, and text processi

    Implements methodologies for red-teaming, constitutional AI, and reward model training to ensure model safety.

    Pythonagentsaiai-agents
    Vezi pe GitHub↗33,575
  • meta-llama/llama3Avatar meta-llama

    meta-llama/llama3

    29,254Vezi pe GitHub↗

    Llama 3 is a collection of pretrained, autoregressive transformer-based models designed for natural language generation, reasoning, and complex instruction following. It functions as a generative AI framework that provides the infrastructure for managing model weights, executing neural network inference, and handling computational workloads across diverse knowledge domains. The project distinguishes itself through an integrated AI safety toolkit that employs secondary classification filtering to inspect inputs and outputs, ensuring adherence to usage compliance and safety standards. It suppor

    Implements secondary classification filtering to inspect inputs and outputs, ensuring adherence to safety and usage compliance standards.

    Python
    Vezi pe GitHub↗29,254
  • elder-plinius/l1b3rt4sAvatar elder-plinius

    elder-plinius/L1B3RT4S

    20,033Vezi pe GitHub↗

    L1B3RT4S is an adversarial machine learning toolkit designed for red teaming and evaluating the robustness of large language models. It provides a research framework for investigating how safety alignment mechanisms and content moderation systems respond to sophisticated input strategies. The project focuses on identifying vulnerabilities in model guardrails by employing techniques such as adversarial narrative framing, dynamic context injection, and latent space steering. It utilizes multi-agent prompt decomposition and recursive text transformation to analyze how structural changes to input

    Investigates the robustness of alignment mechanisms by testing how various prompt engineering techniques influence model output and safety filter behavior.

    1337adversarial-attacksai
    Vezi pe GitHub↗20,033
  • nvidia-nemo/nemoAvatar NVIDIA-NeMo

    NVIDIA-NeMo/NeMo

    17,389Vezi pe GitHub↗

    NeMo is a comprehensive framework designed for the development, training, and deployment of large-scale conversational and generative artificial intelligence models. It provides an integrated platform for building multimodal systems, encompassing speech processing, language modeling, and reinforcement learning alignment. The framework is built to handle the entire lifecycle of AI development, from data curation and model pretraining to production-ready service deployment. The platform distinguishes itself through advanced distributed training capabilities, including tensor and pipeline parall

    Implements programmable guardrails to enforce safety policies and topical boundaries on model outputs.

    Pythonasrdeeplearninggenerative-ai
    Vezi pe GitHub↗17,389
  • googlecloudplatform/generative-aiAvatar GoogleCloudPlatform

    GoogleCloudPlatform/generative-ai

    12,700Vezi pe GitHub↗

    This project is a development platform for managing the lifecycle of generative artificial intelligence models. It provides a unified environment for accessing, fine-tuning, and deploying large language models, serving as an orchestrator that handles the integration of diverse models into custom applications. The platform distinguishes itself by offering a managed infrastructure for hosting and scaling models, which removes the requirement for manual server maintenance or configuration. It includes integrated tools for supervised fine-tuning and vector embedding optimization, allowing for the

    Apply content filtering and responsible guidelines to monitor and restrict model outputs, ensuring that all generated responses adhere to established safety policies and ethical usage standards.

    Jupyter Notebookagentsgcpgemini
    Vezi pe GitHub↗12,700
  • wandb/wandbAvatar wandb

    wandb/wandb

    10,844Vezi pe GitHub↗

    Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow orchestration. It provides a comprehensive suite of tools for logging, visualizing, and versioning training metrics, model artifacts, and hyperparameter sweeps to ensure reproducibility across development cycles. The platform also functions as an observability tool for large language model applications, enabling the tracing of execution steps, token usage, and reasoning processes. The project distinguishes itself through its event-driven automation capabilities, which allow users

    Evaluates model alignment, bias, and safety to ensure responsible behavior.

    Pythonaicollaborationdata-science
    Vezi pe GitHub↗10,844
  • ymcui/chinese-llama-alpaca-2Avatar ymcui

    ymcui/Chinese-LLaMA-Alpaca-2

    7,136Vezi pe GitHub↗

    This project provides a Chinese large language model based on the LLaMA architecture. It is an instruction-tuned model optimized for natural language processing and multi-turn conversations in Chinese. The system includes a framework for parameter-efficient fine-tuning using low-rank adaptation and quantization to reduce memory requirements. It also implements retrieval augmented generation for local document question answering and supports long-context processing for sequences up to 64K tokens. The project covers a broad set of capabilities including supervised instruction tuning, reinforce

    Implements safety and alignment frameworks to ensure model outputs adhere to ethical guidelines.

    Python64kalpacaalpaca-2
    Vezi pe GitHub↗7,136
  • google/gemma_pytorchAvatar google

    google/gemma_pytorch

    5,697Vezi pe GitHub↗

    The official PyTorch implementation of Google's Gemma models

    Implements a safety classifier that evaluates text against policy-defined categories to detect violations.

    Pythongemmagooglepytorch
    Vezi pe GitHub↗5,697
  • dontriskit/awesome-ai-system-promptsAvatar dontriskit

    dontriskit/awesome-ai-system-prompts

    5,206Vezi pe GitHub↗

    This project is a comprehensive library of structured system prompts and configuration templates designed to define the behavior, persona, and operational boundaries of autonomous artificial intelligence agents. It serves as a framework for prompt engineering, providing modular instructions that help models parse complex tasks, maintain consistent interaction tones, and adhere to specific domain constraints. The repository distinguishes itself by offering specialized configurations for agent safety and security, including protocols to prevent prompt injection and unauthorized data access. It

    Provides frameworks for defining safety boundaries and standardized refusal protocols for AI interactions.

    TypeScript
    Vezi pe GitHub↗5,206
  • llm-attacks/llm-attacksAvatar llm-attacks

    llm-attacks/llm-attacks

    4,509Vezi pe GitHub↗

    This repository provides tools and methodologies for studying adversarial attacks on large language models. It focuses on understanding how carefully crafted inputs can manipulate or bypass the safety mechanisms of LLMs, enabling researchers to probe model vulnerabilities and improve their robustness. The project covers techniques for generating adversarial prompts, evaluating model responses under attack conditions, and analyzing the effectiveness of different attack strategies.

    Provides a framework for identifying alignment failures to ensure models do not generate harmful content.

    Python
    Vezi pe GitHub↗4,509
  • zjunlp/easyeditAvatar zjunlp

    zjunlp/EasyEdit

    2,718Vezi pe GitHub↗

    EasyEdit is a framework and toolkit designed for updating, inserting, or erasing specific factual information within large language models without requiring full retraining. It functions as a parameter modifier and knowledge editing system capable of performing targeted weight updates across diverse model architectures. The project distinguishes itself by supporting both text-based and multimodal model editing, allowing for knowledge updates across image and text modalities. It provides utilities for model steering to adjust personality and reasoning patterns in real time via activation inter

    Locates and edits specific model regions to remove toxic behaviors and neutralize harmful outputs.

    Jupyter Notebookartificial-intelligencebaichuanchatgpt
    Vezi pe GitHub↗2,718
  1. Home
  2. Artificial Intelligence & ML
  3. Safety and Alignment Frameworks

Explorează sub-etichetele

  • Content Safety ClassifiersModules that evaluate text against predefined safety policies and flag policy-violating content. **Distinct from Safety and Alignment Frameworks:** Distinct from Safety and Alignment Frameworks: focuses on the classification step itself rather than the broader governance and filtering layer.