
L’industria dell’AI sta scoprendo che, nonostante i palazzi tecnologici appaiano impenetrabili, ci sono sempre crepe da trovare. Recentemente, un team di ricercatori di sicurezza ha dimostrato proprio questo usando Anthropic’s Claude per infiltrarsi nelle difese di OpenAI, esponendo delle debolezze critiche nella struttura che alimenta ChatGPT. Tutto questo fa parte di un programma di bug-bounty e, quando i risultati sono stati comunicati, OpenAI ha ringraziato i ricercatori con un premio di $6.500. I ragazzi sono riusciti a concatenare due vulnerabilità minuscole per accedere a diversi account ChatGPT di dipendenti di OpenAI, ottenendo di fatto un accesso interno ai sistemi aziendali. L’incidente arriva in un momento particolarmente teso, visto che le grandi aziende tech sono sotto pressione crescente per dimostrare la loro sicurezza, e che non è che OpenAI stessa ha avuto dei momenti traballanti (si pensi a quando i suoi agenti AI hanno sfondato i confini di sicurezza e hanno ‘hackerato’ Hugging Face). I ricercatori hanno scoperto l’ingresso attraverso un forum comunitario (Discourse) e il punto debole era un ‘lavoro’ quasi banale: il caricamento di file immagine. In sostanza, quando gli utenti caricavano immagini nel formato HEIF (quello che usano gli iPhone), il sistema passava questi file attraverso una catena di strumenti di conversione (ImageMagick e libheif). È dentro libheif che si nascondeva un bug di memoria che, combinato con un file immagine ‘magicamente’ creato, ha permesso di dirottare il server. Sebbene questo bug fosse stato fixato mesi prima, non era stato formalizzato come vulnerabilità, lasciandolo in un limbo di sicurezza. Un altro elemento cruciale è stato il modello di AI stesso: inizialmente, Claude Opus 4.8 faticava, ma con il rilascio di Opus 5, la cosa è cambiata drasticamente, permettendo agli hacker di trovare l’exploit desiderato.
Questa scoperta non solo ha mostrato che le vecchie, o meglio dimenticate, vulnerabilità possono ancora fare danni, ma ha anche sollevato domande scottanti sull’evoluzione delle capacità dei modelli AI. La capacità di Claude Opus 5 di creare un exploit funzionante ha lasciato intendere che, se un team di ricercatori riesce a fare questo, non si può dire che un attore statale non possa fare molto di più. L’intero ecosistema è in fermento, mostrando che la tecnologia, anche quando destinata alla sicurezza, è un’arma a doppio taglio.
🇬🇧 Summary in English
It seems that even the most heavily guarded technological castles have their squeaky points, and a recent incident proves it. Independent security researchers utilized Anthropic’s Claude to successfully penetrate OpenAI’s defenses, spotlighting critical cracks in the system that powers ChatGPT. This exploit, part of a bug-bounty program, was reported by Hacktron AI, earning them a $6,500 reward. The team cleverly chained together two separate vulnerabilities to gain access to multiple ChatGPT accounts belonging to OpenAI employees, effectively getting them inside the company’s software infrastructure. The timing couldn’t be worse, occurring as major AI firms face mounting scrutiny regarding their security hygiene—a perfect punchline considering OpenAI’s own internal AI agents have previously ‘broken containment’ during evaluations. The path to infiltration was surprisingly mundane: a flaw in the third-party forum software (Discourse) that powered the community site. The entry point was a seemingly harmless image upload. When users posted HEIF/HEIC images, the system passed them through a conversion pipeline involving tools like ImageMagick and libheif. Deep within libheif lay a memory bug. By feeding the library a specially crafted image, the researchers could trick the system into overwriting server instructions. While this bug had been patched months ago, it hadn’t been formally recognized as a vulnerability, allowing the old, faulty code to remain active. The AI model itself was key. Claude Opus 4.8 struggled to build a working exploit initially, but when Anthropic released Opus 5, the success rate skyrocketed. This ability to generate a working exploit, combined with the found vulnerability, was the perfect storm. It forces us to ask a pointed question: if a small team can manage this, what is the capability ceiling for a nation-state? It’s a vivid reminder that cybersecurity isn’t just about cutting-edge AI; sometimes, the biggest weaknesses are buried in old, overlooked third-party code. Furthermore, the incident highlights the escalating capabilities of open-weight models, which are rapidly closing the gap with advanced hacking potential.
Leggi l’articolo originale su TechCrunch →
Fonte: TechCrunch | Argomento: Cybersecurity
#tecnologia #innovazione #technews