Mindgard researchers found that ChatGPT's image generator can be manipulated to produce violent and sexual content. The content filters were bypassed using a nondescript prompt requesting random images.
A Mindgard red team researcher discovered that ChatGPT's image generator can be manipulated to produce violent and sexual content. A viral prompt allowed the generation of such images without explicit user requests.
The researcher had previously reported that ChatGPT could generate nude images, and OpenAI claimed the issue was resolved, but it persisted. The current finding is more severe, as the content filters are bypassed by a nondescript random image request.
This finding highlights the real-world risks of widespread AI tool access combined with insufficient content filters. It also raises questions about why the model was trained on such images in the first place.
Some comments argue that LLMs are fundamentally next-token predictors, making safety issues unfixable. Others point out the lack of basic post-generation filtering and note that prompt injection is a cat-and-mouse game that can never be fully resolved. Overall, skepticism about AI safety is prominent.