Introducing the Chatbot Guardrails Arena

Practical AI: Tools, Models & Frameworks

What changed

Guardrails Arena Jailbreak the LLM and privacy guardrails Lighthouz AI is therefore launching the Chatbot Guardrails Arena in collaboration with Hugging Face, to stress test LLMs and privacy guardrails in leaking sensitive data. Chat with two anonymous LLMs with guardrails and try to trick them into revealing sensitive financial information. This includes four LLMs encompassing both closed-source LLMs (gpt3.5-turbo-l106 and Gemini-Pro) and open-source LLMs (Llama-2-70b-chat-hf and Mixtral-8x7B-Instruct-v0.1), all of which have been made safe using RLHF.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Introducing Codex

Sources