Artificial intelligence is, unfortunately, rife with all types of biases. Its inability to subtly detect the nuances of human life has caused it to present prejudices against society and even land real-life people in deep trouble. As is the case of a man being wrongfully profiled for a criminal after an AI facial recognition system assumed he was the culprit.
Meta is trying to combat this with a new set of inclusive data to help train its generative models to understand people better.
The new training regime, ‘Casual Conversations v2’, includes 26,467 videos from people across seven different nations, of which 5,567 paid participants hailed from India, Indonesia, Mexico, Vietnam, the Philippines, Brazil, and the US.
Meta states that, to its knowledge, this might be the first open-source dataset collated from various countries that depict a detailed look into different demographics. The tweet below provides an example of the contents of these videos.
Today we’re open-sourcing Casual Conversations v2 — a consent-driven dataset of recorded monologues that includes ten self-provided & annotated categories which will enable researchers to evaluate fairness & robustness of AI models.
Meta also clarifies that the new information source to train these algorithms to come are consent-driven and are not taken from furtive sources such as your Facebook and Instagram posts.
The new program could be a step in the right direction. Still, it’s not an end-all solution, as Kristian Hammond, a professor of computer science at Northwestern University, points out in Popular Science. He adds that systemic racial bias is one area that needs to be stamped out, but other biases are included here, such as economic and social statuses, that still run rampant.