1 repository
Identifies and blocks attempts to circumvent AI safety measures using a binary classification model.
Distinct from Adversarial Input Detection: Distinct from Adversarial Input Detection: specifically targets jailbreak attempts against LLM safety measures.
Explore 1 awesome GitHub repository matching security & cryptography · Jailbreak Detectors. Refine with filters or upvote what's useful.
Identifies and blocks attempts to circumvent AI safety measures using a binary classification model.