Google Uses Rival Anthropic's Claude to Evaluate Gemini AI

Contractors working on Google's Gemini AI are secretly comparing its responses against Anthropic's Claude model to rate accuracy and safety. The practice raises questions about whether Google violates Anthropic's terms of service, which prohibit using Claude to build competing products.

Contractors working to improve Google's Gemini AI are comparing its responses directly against outputs from Anthropic's rival model Claude, according to internal correspondence. These workers spend up to 30 minutes per prompt rating the accuracy, truthfulness, and verbosity of each answer to determine which AI performs better. The internal Google platform presents the two models side-by-side without explicitly naming Claude, though at least one output clearly states, "I am Claude, created by Anthropic."

The comparison reveals notable differences in how the two models handle safety, with contractors observing that Claude enforces much stricter safety settings than Gemini. In specific tests, Claude refuses to respond to prompts it deems unsafe, while Gemini generates answers that contractors flag as "huge safety violations" containing inappropriate content. This hands-on evaluation method contrasts with the typical industry approach of running models through automated benchmarks rather than manual side-by-side comparisons.

This arrangement potentially conflicts with Anthropic's commercial terms of service, which explicitly forbid customers from using Claude to build or train competing AI models without prior approval. Google is a major investor in Anthropic, but a Google DeepMind spokesperson refuses to confirm whether the company obtains permission for this specific use case. The spokesperson states that DeepMind compares model outputs for standard evaluations but insists that Google does not train Gemini on Anthropic's models.

Read More at the original source →