Anthropic claims its new AI chatbot models beat OpenAI’s GPT-4

11:50 AM PST • March 4, 2024

Image Credits: Anthropic

AI startup Anthropic, backed by Google and hundreds of millions in venture capital (and perhaps soon hundreds of millions more), today announced the latest version of its GenAI tech, Claude. And the company claims that the AI chatbot beats OpenAI’s GPT-4 in terms of performance.

Claude 3, as Anthropic’s new GenAI is called, is a family of models — Claude 3 Haiku, Claude 3 Sonnet and Claude 3 Opus, Opus being the most powerful. All show “increased capabilities” in analysis and forecasting, Anthropic claims, as well as enhanced performance on specific benchmarks versus models like ChatGPT and GPT-4 and Google’s Gemini 1.0 Ultra (but not Gemini 1.5 Pro).

Notably, Claude 3 is Anthropic’s first multimodal GenAI, meaning that it can analyze text as well as images — similar to some flavors of GPT-4 and Gemini. Claude 3 can process photos, charts, graphs and technical diagrams, drawing from PDFs, slideshows and other document types.

In a step one better than some GenAI rivals, Claude 3 can analyze multiple images in a single request (up to a maximum of 20). This allows it to compare and contrast images, notes Anthropic.

But there are limits to Claude 3’s image processing.

Anthropic has disabled the models from identifying people — no doubt wary of the ethical and legal implications. And the company admits that Claude 3 is prone to making mistakes with “low-quality” images (under 200 pixels) and struggles with tasks involving spatial reasoning (e.g. reading an analog clock face) and object counting (Claude 3 can’t give exact counts of objects in images).

Anthropic Claude 3 — **Image Credits:** Anthropic

Claude 3 also won’t generate artwork. The models are strictly image-analyzing — at least for now.

Whether fielding text or images, Anthropic says that customers can generally expect Claude 3 to better follow multi-step instructions, produce structured output in formats like JSON and converse in languages other than English compared to its predecessors. Claude 3 should also refuse to answer questions less often thanks to a “more nuanced understanding of requests,” Anthropic says. And soon, the models will cite the source of their answers to questions so users can verify them.

“Claude 3 tends to generate more expressive and engaging responses,” Anthropic writes in a support article. “[It’s] easier to prompt and steer compared to our legacy models. Users should find that they can achieve the desired results with shorter and more concise prompts.”

Some of those improvements stem from Claude 3’s expanded context.

A model’s context, or context window, refers to input data (e.g. text) that the model considers before generating output. Models with small context windows tend to “forget” the content of even very recent conversations, leading them to veer off topic — often in problematic ways. As an added upside, large-context models can better grasp the narrative flow of data they take in and generate more contextually rich responses (hypothetically, at least).

Anthropic says that Claude 3 will initially support a 200,000-token context window, equivalent to about 150,000 words, with select customers getting up a 1-milion-token context window (~700,000 words). That’s on par with Google’s newest GenAI model, the above-mentioned Gemini 1.5 Pro, which also offers up to a million-token context window.

Now, just because Claude 3 is an upgrade over what came before it doesn’t mean it’s perfect.

In a technical whitepaper, Anthropic admits that Claude 3 isn’t immune from the issues plaguing other GenAI models, namely bias and hallucinations (i.e. making stuff up). Unlike some GenAI models, Claude 3 can’t search the web; the models can only answer questions using data from before August 2023. And while Claude is multilingual, it’s not as fluent in certain “low-resource” languages versus English.

But Anthropic is promising frequent updates to Claude 3 in the months to come.

“We don’t believe that model intelligence is anywhere near its limits, and we plan to release [enhancements] to the Claude 3 model family over the next few months,” the company writes in a blog post.

Opus and Sonnet are available now on the web and via Anthropic’s dev console and API, Amazon’s Bedrock platform and Google’s Vertex AI. Haiku will follow later this year.

Here’s the pricing breakdown:

Opus: $15 per million input tokens, $75 per million output tokens
Sonnet: $3 per million input tokens, $15 per million output tokens
Haiku: $0.25 per million input tokens, $1.25 per million output tokens

So that’s Claude 3. But what’s the 30,000-foot view of all this?

Well, as we’ve reported previously, Anthropic’s ambition is to create a next-gen algorithm for “AI self-teaching.” Such an algorithm could be used to build virtual assistants that can answer emails, perform research and generate art, books and more — some of which we’ve already gotten a taste of with the likes of GPT-4 and other large language models.

Anthropic hints at this in the aforementioned blog post, saying that it plans to add features to Claude 3 that enhance its out-of-the-gate capabilities by allowing Claude to interact with other systems, code “interactively” and deliver “advanced agentic capabilities.”

That last bit calls to mind OpenAI’s reported ambitions to build a software agent to automate complex tasks, like transferring data from a document to a spreadsheet or automatically filling out expense reports and entering them in accounting software. OpenAI already offers an API that allows developers to build “agent-like experiences” into their apps, and Anthropic, it seems, is intent on delivering functionality that’s comparable.

Could we see an image generator from Anthropic next? It’d surprise me, frankly. Image generators are the subject of much controversy these days, mainly for copyright- and bias-related reasons. Google was recently forced to disable its image generator after it injected diversity into pictures with a farcical disregard for historical context. And a number of image generator vendors are in legal battles with artists who accuse them of profiting off of their work by training GenAI on that work without providing compensation or even credit.

I’m curious to see the evolution of Anthropic’s technique for training GenAI, “constitutional AI,” which the company claims makes the behavior of its GenAI easier to understand, more predictable and simpler to adjust as needed. Constitutional AI aims to provide a way to align AI with human intentions, having models respond to questions and perform tasks using a simple set of guiding principles. For example, for Claude 3, Anthropic said that it added a principle — informed by crowdsourced feedback — that instructs the models to be understanding of and accessible to people with disabilities.

Whatever Anthropic’s endgame, it’s in it for the long haul. According to a pitch deck leaked in May of last year, the company aims to raise as much as $5 billion over the next 12 months or so — which might just be the baseline it needs to remain competitive with OpenAI. (Training models isn’t cheap, after all.) It’s well on its way, with $2 billion and $4 billion in committed capital and pledges from Google and Amazon, respectively, and well over a billion combined from other backers.

More TechCrunch

Meta’s Oversight Board takes its first Threads case

Sarah Perez

23 mins ago

Meta’s Oversight Board has now extended its scope to include the company’s newest platform, Instagram Threads. Designed as an independent appeals board that hears cases and then makes precedent-setting content…

Meta’s Oversight Board takes its first Threads case

Venture

SeekOut, a recruiting startup last valued at $1.2 billion, lays off 30% of its workforce

Marina Temkin

Mary Ann Azevedo

1 hour ago

The company says it’s refocusing and prioritizing fewer initiatives that will have the biggest impact on customers and add value to the business.

SeekOut, a recruiting startup last valued at $1.2 billion, lays off 30% of its workforce

Transportation

UK’s autonomous vehicle legislation becomes law, paving the way for first driverless cars by 2026

Paul Sawers

2 hours ago

The U.K.’s self-proclaimed “world-leading” regulations for self-driving cars are now official, after the Automated Vehicles (AV) Act received royal assent — the final rubber stamp any legislation must go through…

UK’s autonomous vehicle legislation becomes law, paving the way for first driverless cars by 2026

ChatGPT: Everything you need to know about the AI-powered chatbot

Alyssa Stringer

Kyle Wiggers

Cody Corrall

2 hours ago

ChatGPT, OpenAI’s text-generating AI chatbot, has taken the world by storm. What started as a tool to hyper-charge productivity through writing essays and code with short text prompts has evolved…

ChatGPT: Everything you need to know about the AI-powered chatbot

Fintech

Fintech lender SoLo Funds is being sued again by the government over its lending practices

Christine Hall

3 hours ago

SoLo Funds CEO Travis Holoway: “Regulators seem driven by press releases when they should be motivated by true consumer protection and empowering equitable solutions.”

Fintech lender SoLo Funds is being sued again by the government over its lending practices

Enterprise

Rollup wants to be the hardware engineer’s workhorse

Aria Alamalhodaei

3 hours ago

Hard tech startups generate a lot of buzz, but there’s a growing cohort of companies building digital tools squarely focused on making hard tech development faster, more efficient and —…

Rollup wants to be the hardware engineer’s workhorse

Startups

Disrupt Audience Choice vote closes Friday

TechCrunch Events

3 hours ago

TechCrunch Disrupt 2024 is not just about groundbreaking innovations, insightful panels, and visionary speakers — it’s also about listening to YOU, the audience, and what you feel is top of…

Disrupt Audience Choice vote closes Friday

Apps

Google is launching a new Android feature to drive users back into their installed apps

Sarah Perez

3 hours ago

Google says the new SDK would help Google expand on its core mission of connecting the right audience to the right content at the right time.

Google is launching a new Android feature to drive users back into their installed apps

Privacy

Jolla debuts privacy-focused AI hardware

Natasha Lomas

3 hours ago

Jolla has taken the official wraps off the first version of its personal server-based AI assistant in the making. The reborn startup is building a privacy-focused AI device — aka…

Jolla debuts privacy-focused AI hardware

OpenAI to remove ChatGPT’s Scarlett Johansson-like voice

Aisha Malik

4 hours ago

OpenAI is removing one of the voices used by ChatGPT after users found that it sounded similar to Scarlett Johansson, the company announced on Monday. The voice, called Sky, is…

OpenAI to remove ChatGPT’s Scarlett Johansson-like voice

ChatGPT’s mobile app revenue saw its biggest spike yet following GPT-4o launch

Sarah Perez

4 hours ago

The ChatGPT mobile app’s net revenue first jumped 22% on the day of the GPT-4o launch and continued to grow in the following days.

ChatGPT’s mobile app revenue saw its biggest spike yet following GPT-4o launch

Social

Bumble buys community building app Geneva to expand further into friendships

Paul Sawers

6 hours ago

Dating app maker Bumble has acquired Geneva, an online platform built around forming real-world groups and clubs. The company said that the deal is designed to help it expand its…

Bumble buys community building app Geneva to expand further into friendships

Security

CyberArk snaps up Venafi for $1.54B to ramp up in machine-to-machine security

Ingrid Lunden

7 hours ago

CyberArk — one of the army of larger security companies founded out of Israel — is acquiring Venafi, a specialist in machine identity, for $1.54 billion.

CyberArk snaps up Venafi for $1.54B to ramp up in machine-to-machine security

Venture

OpenseedVC, which backs operators in Africa and Europe starting their companies, reaches first close of $10M fund

Tage Kene-Okafor

11 hours ago

Founder-market fit is one of the most crucial factors in a startup’s success, and operators (someone involved in the day-to-day operations of a startup) turned founders have an almost unfair advantage…

OpenseedVC, which backs operators in Africa and Europe starting their companies, reaches first close of $10M fund

Startups

Pine Labs gets Singapore court approval to shift base to India

Manish Singh

13 hours ago

A Singapore High Court has effectively approved Pine Labs’ request to shift its operations to India.

Pine Labs gets Singapore court approval to shift base to India

Government & Policy

UK opens office in San Francisco to tackle AI risk

Ingrid Lunden

19 hours ago

The AI Safety Institute, a U.K. body that aims to assess and address risks in AI platforms, has said it will open a second location in San Francisco.

UK opens office in San Francisco to tackle AI risk

Enterprise

Why companies are turning to internal hackathons

Ron Miller

1 day ago

Companies are always looking for an edge, and searching for ways to encourage their employees to innovate. One way to do that is by running an internal hackathon around a…

Why companies are turning to internal hackathons

Featured Article

I’m rooting for Melinda French Gates to fix tech’s broken ‘brilliant jerk’ culture

Women in tech still face a shocking level of mistreatment at work. Melinda French Gates is one of the few working to change that.

Julie Bort

1 day ago

I’m rooting for Melinda French Gates to fix tech’s broken ‘brilliant jerk’ culture

Space

Blue Origin successfully launches its first crewed mission since 2022

Anthony Ha

1 day ago

Blue Origin has successfully completed its NS-25 mission, resuming crewed flights for the first time in nearly two years. The mission brought six tourist crew members to the edge of…

Blue Origin successfully launches its first crewed mission since 2022

Hollywood agency CAA aims to help stars manage their own AI likenesses

Lauren Forristal

1 day ago

Creative Artists Agency (CAA), one of the top entertainment and sports talent agencies, is hoping to be at the forefront of AI protection services for celebrities in Hollywood. With many…

Hollywood agency CAA aims to help stars manage their own AI likenesses

Commerce

Expedia says two execs dismissed after ‘violation of company policy’

Anthony Ha

1 day ago

Expedia says Rathi Murthy and Sreenivas Rachamadugu, respectively its CTO and senior vice president of core services product & engineering, are no longer employed at the travel booking company. In…

Expedia says two execs dismissed after ‘violation of company policy’

Social

OpenAI and Google lay out their competing AI visions

Cody Corrall

2 days ago

Welcome back to TechCrunch’s Week in Review. This week had two major events from OpenAI and Google. OpenAI’s spring update event saw the reveal of its new model, GPT-4o, which…

OpenAI and Google lay out their competing AI visions

Startups

With AI startups booming, nap pods and Silicon Valley hustle culture are back

Julie Bort

2 days ago

When Jeffrey Wang posted to X asking if anyone wanted to go in on an order of fancy-but-affordable office nap pods, he didn’t expect the post to go viral.

With AI startups booming, nap pods and Silicon Valley hustle culture are back

OpenAI created a team to control ‘superintelligent’ AI — then let it wither, source says

Kyle Wiggers

2 days ago

OpenAI’s Superalignment team, responsible for developing ways to govern and steer “superintelligent” AI systems, was promised 20% of the company’s compute resources, according to a person from that team. But…

OpenAI created a team to control ‘superintelligent’ AI — then let it wither, source says

Transportation

VCs and the military are fueling self-driving startups that don’t need roads

Rebecca Bellan

2 days ago

A new crop of early-stage startups — along with some recent VC investments — illustrates a niche emerging in the autonomous vehicle technology sector. Unlike the companies bringing robotaxis to…

VCs and the military are fueling self-driving startups that don’t need roads

Enterprise

Deal Dive: Sagetap looks to bring enterprise software sales into the 21st century

Rebecca Szkutak

2 days ago

When the founders of Sagetap, Sahil Khanna and Kevin Hughes, started working at early-stage enterprise software startups, they were surprised to find that the companies they worked at were trying…

Deal Dive: Sagetap looks to bring enterprise software sales into the 21st century

This Week in AI: OpenAI moves away from safety

Kyle Wiggers

Devin Coldewey

2 days ago

Keeping up with an industry as fast-moving as AI is a tall order. So until an AI can do it for you, here’s a handy roundup of recent stories in the world…

This Week in AI: OpenAI moves away from safety

Apps

Adobe comes after indie game emulator Delta for copying its logo

Sarah Perez

3 days ago

After Apple loosened its App Store guidelines to permit game emulators, the retro game emulator Delta — an app 10 years in the making — hit the top of the…

Apps

Meta’s latest experiment borrows from BeReal’s and Snapchat’s core ideas

Aisha Malik

3 days ago

Meta is once again taking on its competitors by developing a feature that borrows concepts from others — in this case, BeReal and Snapchat. The company is developing a feature…