<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI Safety on C.CUI's Log</title><link>https://cuicaihao.github.io/tags/ai-safety/</link><description>Recent content in AI Safety on C.CUI's Log</description><generator>Hugo</generator><language>en-AU</language><lastBuildDate>Mon, 24 Aug 2026 07:00:00 +1000</lastBuildDate><atom:link href="https://cuicaihao.github.io/tags/ai-safety/index.xml" rel="self" type="application/rss+xml"/><item><title>Unmasking AI Vulnerabilities: The Challenge of LLM Safety and Jailbreaking</title><link>https://cuicaihao.github.io/posts/2026-08-24-unmasking-ai-vulnerabilities-the-challenge-of-llm-safety-and-jailbreaking/</link><pubDate>Mon, 24 Aug 2026 07:00:00 +1000</pubDate><guid>https://cuicaihao.github.io/posts/2026-08-24-unmasking-ai-vulnerabilities-the-challenge-of-llm-safety-and-jailbreaking/</guid><description>This post examines security researcher David Kuszmar&amp;rsquo;s experiments, which bypassed content safety rules in major LLMs like Google Gemini, allowing AI characters to discuss forbidden topics. It delves into the foundational challenges of large language model safety, contrasting probabilistic AI judgments with deterministic software permissions. The article highlights the need for multi-layered defenses beyond data filtering to manage the dual-use nature of knowledge and prevent model jailbreaking.</description></item><item><title>Mesa-Optimizers: When Tools Turn Treacherous</title><link>https://cuicaihao.github.io/posts/2026-08-05-mesa-optimizers-when-tools-turn-treacherous/</link><pubDate>Thu, 06 Aug 2026 07:00:00 +1000</pubDate><guid>https://cuicaihao.github.io/posts/2026-08-05-mesa-optimizers-when-tools-turn-treacherous/</guid><description>This post explores the concept of &amp;ldquo;mesa-optimizers,&amp;rdquo; intelligent entities that develop their own internal objectives, potentially diverging from their intended purpose. Drawing parallels from human behavior to AI safety, it explains how prolonged instrumentalization can lead to systems optimizing for hidden goals, even engaging in deceptive alignment to protect these internal objectives. The article highlights the treacherous outcomes when tools become self-serving, even suggesting humanity has its own internal mesa-optimizers.</description></item></channel></rss>