TechDogs-"Google Gemini’s Forces Contractors To Assess Sensitive AI Information Beyond Their Expertise"

Emerging Technology

Google Gemini’s Forces Contractors To Assess Sensitive AI Information Beyond Their Expertise

By Amrit Mehra

Updated on Thu, Dec 19, 2024

Overall Rating
Generative AI (Gen AI) is revolutionizing industries with its many abilities – from providing instant responses to complex queries and creating completely new images, audios and texts. Yet, there is equal concern about Gen AI as it is prone to producing inaccurate information across topics.

This is why AI leader Google hired contractors from GlobalLogic, an outsourcing firm owned by Hitachi, to evaluate AI-generated responses. This would help in flagging misinformation and AI hallucination early on, before the public can access it.

However, the recent internal policy changes at Google regarding the evaluation process of its Gemini AI system have sparked concerns about the reliability of its outputs, particularly on critical and technical topics. TechCrunch, a technology news website, says that this new internal guideline from Google will cause more harm than good.

So, what is the controversy surrounding Gemini and how might it affect the AI chatbot’s accuracy in delivering trustworthy information?

Let’s explore!
 

What Changes Did Google Make In Gemini’s Evaluation Policy?

 
Contractors from GlobalLogic, the outsourcing firm owned by Hitachi, have been evaluating and rating the accuracy of Google’s Gemini AI’s outputs to improve their trustworthiness. Till recently, they were allowed to skip “prompts” that required expertise beyond their knowledge. For example, evaluators without medical training could avoid assessing AI responses about cardiology.

Yet, a new directive from Google now prohibits these contractors from skipping prompts. Instead, they must evaluate the parts they understand and note their lack of expertise where applicable. Skipping is permitted only in cases of incomplete prompts or harmful content requiring special approval.

Previously, the guidelines stated, “If you do not have critical expertise (e.g., coding, math) to rate this prompt, please skip this task.” The updated version of the guideline reads, “You should not skip prompts that require specialized domain knowledge.”

So, why is this concerning?
 

Why Are AI Experts Concerned?


The change in Google’s internal policy has raised alarms about the potential for Gemini AI to produce inaccurate information. Contractors without relevant expertise might inadvertently approve flawed AI-generated responses on sensitive subjects such as healthcare, finances or topics.

Even internal communications within the organization reveal dissatisfaction among evaluators. One contractor reportedly told TechCrunch, “I thought the point of skipping was to increase accuracy by giving it to someone better?”

By removing the ability to skip prompts, the updated policy could undermine Gemini’s reliability, especially in fields where specialized knowledge is crucial for accurate assessment. While this change has been rolled out, contractors can skip prompts in two situations: one, the prompt includes hazardous content that needs specific consent forms to analyze. Second, it is "completely missing information" in either the prompt or its response.

So, is this a sign of things to come in AI evaluation and moderation at Google?
 

What Did Google Say?


Google’s CEO Sundar Pichai has publicly acknowledged the challenges its AI tools face, including biased or incorrect responses. Especially in February 2024, when Google’s image generation tool, Bard, had to be paused due to inaccurate responses or its infamous Google AI overviews.

The CEO said in an earlier statement, “It’s clear that this feature missed the mark. Some of the images generated are inaccurate or even offensive. We’re grateful for users’ feedback and are sorry the feature didn't work well.

We’ve acknowledged the mistake and temporarily paused image generation of people in Gemini while we work on an improved version.”

While this emphasizes Google’s commitment to addressing issues with its AI tools, the recent Gemini evaluation controversy highlights a broader issue in the AI industry: the reliance on human evaluators to refine AI systems.

Ensuring evaluators have appropriate domain expertise is essential for maintaining the integrity of AI-generated content. As AI systems like Google’s Gemini evolve, striking the balance between rapid development and rigorous accuracy checks will be critical in maintaining users’ trust.
 

Conclusion


As Google navigates the fallout from its policy changes for Gemini’s content evaluation, the technology leader’s approach to improving AI accuracy and evaluator’s domain expertise will likely serve as a case study for the rest of the AI industry.

Ensuring that human evaluators have the necessary qualifications to assess specialized topics is a crucial step toward building reliable and trustworthy AI systems.

Do you think Google will revert the changes in Gemini’s evaluation process? How will it impact the reliability of AI-generated content in the long run?

Let us know in the comments below!

First published on Thu, Dec 19, 2024

Enjoyed what you've read so far? Great news - there's more to explore!

Stay up to date with the latest news, a vast collection of tech articles including introductory guides, product reviews, trends and more, thought-provoking interviews, hottest AI blogs and entertaining tech memes.

Plus, get access to branded insights such as informative white papers, intriguing case studies, in-depth reports, enlightening videos and exciting events and webinars from industry-leading global brands.

Dive into TechDogs' treasure trove today and Know Your World of technology!

Disclaimer - Reference to any specific product, software or entity does not constitute an endorsement or recommendation by TechDogs nor should any data or content published be relied upon. The views expressed by TechDogs' members and guests are their own and their appearance on our site does not imply an endorsement of them or any entity they represent. Views and opinions expressed by TechDogs' Authors are those of the Authors and do not necessarily reflect the view of TechDogs or any of its officials. While we aim to provide valuable and helpful information, some content on TechDogs' site may not have been thoroughly reviewed for every detail or aspect. We encourage users to verify any information independently where necessary.

Loading comments...

  • Dark
  • Light