
Photo 273969814 © Darkworx | Dreamstime.com
Ever coded a masterpiece only to discover a pesky bug lurking beneath the surface? Well, developers using OpenAI’s ChatGPT may soon have a secret weapon: CriticGPT. Rather than judging your Python coding style, this artificial intelligence watchdog is here to sniff out errors in code generated by its predecessor and catch inefficiencies before they slip through the cracks.
Built on the same GPT-4 architecture as ChatGPT, CriticGPT analyzes code and flags up any issues to ensure the final product is polished and functional.
The organization says the results have been promising. “We found that when people get help from CriticGPT to review ChatGPT code they outperform those without help 60% of the time.”
So, how does CriticGPT learn to be such a hawk-eyed critic? OpenAI employed a technique called Reinforcement Learning from Human Feedback (RLHF). It fed the model with a massive dataset of buggy code, including both natural errors produced by ChatGPT and ones deliberately inserted by human trainers. Imitating a human reviewer, the trainers then provided feedback on how to fix these errors. CriticGPT then compared these critiques to hone its own bug-spotting skills.
As AI models like ChatGPT become more sophisticated in reasoning, their errors become subtler and harder for humans to detect. This challenge hinders the RLHF process, where human feedback is crucial for training.
Moving forward, OpenAI is keen on integrating CriticGPT into its RLHF labeling pipeline to offer explicit AI assistance to trainers.
“In order to align AI systems that are increasingly complex, we’ll need better tools,” the company concludes. “We are planning to scale this work further and put it into practice.”
[via Windows Central, MarkTechPost, Money Control, images via various sources]