Startups

Why image recognition is about to transform business

Comment

Ken Weiner

Contributor

Ken Weiner is the CTO of GumGum.

At Facebook’s recent annual developer conference, Marc Zuckerberg outlined the social network’s artificial intelligence (AI) plans to “build systems that are better than people in perception.” He then demonstrated an impressive image recognition technology for the blind that can “see” what’s going on in a picture and explain it out loud.

From programs that help the visually impaired and safety features in cars that detect large animals to auto-organizing untagged photo collections and extracting business insights from socially shared pictures, the benefits of image recognition, or computer vision, are only just beginning to make their way into the world — but they’re doing so with increasing frequency and depth.

It’s busy enough that the upcoming LDV Vision Summit, an annual conference dedicated to all things visual tech, from VR and cameras to medical imaging and content analysis, is already in its third year. “The advancements in computer vision these days are creating tremendous new opportunities in analyzing images that are exponentially impacting every business vertical, from automotive to advertising to augmented reality,” says Evan Nisselson of LDV Capital, which organizes the summit.

As with other forms of AI — natural language procession, bioinformatics, gaming — the field of computer vision has benefited greatly from the expansion of open-source, deep learning technology, user-friendly programming tools and faster and more affordable computing.

Many a headline references deep learning and artificial intelligence as the next big thing, but how exactly do these different tools work, and in what ways are businesses using them to offer image tech to the world? Is Google’s TensorFlow the same thing as Facebook’s DeepFace or Microsoft’s Project Oxford? Not exactly. To help clarify things, here’s a quick breakdown of current image technology tools and how businesses are using them.

Training material: Open data

Thanks to deep learning techniques, a machine learning technique loosely modeled after the human brain, computers can be taught to accurately identify what’s in pictures faster than ever — but they need massive amounts of data to do it.

Enter ImageNet and Pascal VOC. Years in the making, these massive and free-to-anyone databases contain millions of images tagged with keywords about what’s inside the pictures — everything from cats and mountains to pizza and sports activities. These open datasets are the basis for machine learning around images (the only way computers can accurately identify cats in photos is because they have already learned what cats look like by analyzing millions of pictures tagged with the word “cat”).

Best known for its annual visual recognition challenge, ImageNet was launched by computer scientists at Stanford and Princeton in 2009 with 80,000 tagged images. It has since grown to include more than 14 million tagged images, any of which are up for grabs at any time for machine training purposes.

Powered by various universities in the U.K., Pascal VOC has fewer pictures, but each one has richer annotations. This improves the accuracy and breadth of the machine learning and, for some applications, speeds up the overall process, because it allows for the omission of cumbersome computer subtasks.

Now, everyone from Google and Facebook to startups and universities use these open source picture sets to feed their machine learning beasts, but the big technology companies have the advantage of access to millions of user-labeled images from apps such as Google Photos and Facebook. Have you ever wondered why Google and Facebook let you upload so many pictures for free? It’s because those pictures are used to train their deep learning networks to become more accurate.

Building blocks: Open-source software libraries and frameworks

Once you have the data, it’s time to build a machine that can learn from it. Enter open-source software libraries. Freely available, these frameworks serve as starting points for building machine learning systems to service different kinds of computer vision functions, from facial and emotion recognition to medical screening and large obstacle (read: deer) detection in cars. These machine learning systems are then fed pictures from ImageNet and its ilk, proprietary images (aka Google Photos) or other sources (like anonymized, indexed clinical records).

Google TensorFlow is one of the better-known libraries, if only because it was covered widely when selected parts were open sourced late last year. TensorFlow, some of which is still proprietary to Google, is used to develop many of the company’s AI initiatives, from autonomous cars and translation to Google Now and Google Photos.

But TensorFlow is hardly the first — or only — open-source framework. UC Berkeley’s Caffe has been around since 2009, and remains popular because of its ease of customizability and large community of innovators, not to mention heavy use by Pinterest and Yahoo!/Flickr. Even Google turns to Caffe for certain projects such as DeepDream.

Created in 2002, Torch is also popular, owing to its use by Facebook AI Research (FAIR), which open sourced some of its modules in early 2015. Some of these tools are optimized to run on more than one graphics processor or computer to amplify capacity and speed up the deep learning process. Similarly, NVIDIA’s cuDNN is an open-source software library that optimizes a computer’s graphics processing unit (GPU) performance, making machine learning even faster.

These tools, while flexible and robust, require teams of computer vision engineers and hardware, so only companies that want to make computer vision a major part of their product strategy, where they’d want to own the software, need apply.

Ready-to-wear: Hosted APIs

Not every company has the resources, or wants to invest in the resources, to build out a computer vision engineering team. Even if you’ve found the right team, it can be a lot of work to get it just right, which is where hosted API services come in. Carried out in the cloud, these solutions offer menus of out-of-the-box image recognition services that can be easily integrated with an existing app or used to build out a specific feature or an entire business.

Say the Travel Channel needs “landmark detection” to show relevant photos on landing pages for specific landmarks, or eHarmony wants to filter out “unsafe” profile images uploaded by their users. Neither of these companies needs or wants to get into the deep learning image recognition development business, but can still benefit from its capabilities.

Google Cloud Vision, for example, offers a series of image detection services from facial and optical character recognition (text) to landmark and explicit content detection, and charges on a per-photo basis. Microsoft Cognitive Services (née Project Oxford) offers a collection of visual image recognition APIs, including emotion, celebrity and face detection, and charges a specific rate per 1,000 transactions. Meanwhile, startups like Clarifai offer computer vision APIs that help companies organize their content, filter out unsafe user-generated images and videos and make purchasing recommendations based on viewed or taken photos.

Custom computer vision technology

Of course, it doesn’t have to be apples or oranges. Computer vision engineering teams don’t need to be Google-sized, and companies big and small that don’t want to build their own AI systems may still want robust, custom image recognition solutions. If a beauty or cosmetics company wants to find, say, pictures of people with high-volume hair to serve ads about body-minimizing shampoo, it’ll need someone to create a custom algorithm to search for high-volume hair, since that isn’t the first thing that the more commoditized solutions offer out of the box.

Same with logos or car make and model, which are still niche commercial applications that currently aren’t available in the open-source arena. And if a closed dataset isn’t readily available, no matter, because a good percentage of the images shared on social media these days are public, anyway, making for a rich source of images with which to feed the machine learning beast.

Some companies use combinations of open data and open-source frameworks, as long as they have a team of engineers, or they might just use hosted APIs if computer vision is not something on which they are staking their entire business.

And for companies with a wide range of very specific needs, there are custom solutions. No matter how it’s approached, though, it’s clear that image recognition rarely exists in isolation; it’s made stronger by access to more and more pictures, real-time big data, unique applications and speed. The businesses that make the most of these connections are the ones that will be best poised for success.

More TechCrunch

OpenAI’s chatbot ChatGPT has been down for several users across the globe for the last few hours.

ChatGPT is down for several users, OpenAI is working on a fix

Microsoft’s education-focused flavor of its cloud productivity suite, Microsoft 365 Education, is facing investigation in the European Union where privacy rights non-profit noyb has just lodged two complaints with Austria’s…

Microsoft hit with EU privacy complaints over schools’ use of 365 Education suite

Since the shock of Russia’s 2022 invasion of Ukraine, solar energy has been having a moment in Europe. Electricity prices have been going up while the investment required to get…

Samara is accelerating the energy transition in Spain one solar panel at a time

Featured Article

DEI backlash: Stay up-to-date on the latest legal and corporate challenges

It’s clear that this year will be a turning point for DEI.

11 hours ago
DEI backlash: Stay up-to-date on the latest legal and corporate challenges

The keynote will be focused on Apple’s software offerings and the developers that power them, including the latest versions of iOS, iPadOS, macOS, tvOS, visionOS and watchOS.

Watch Apple kick off WWDC 2024 right here

Hello and welcome back to TechCrunch Space. Unfortunately, Boeing’s Starliner launch was delayed yet again, this time due to issues with one of the three redundant computers used by United…

TechCrunch Space: China’s victory

The court ruling said that Fearless Fund’s Strivers Grant likely violates the Civil Rights Act of 1866, which bans the use of race in contracts.

An appeals court rules that VC Fearless Fund cannot issue grants to Black women, but the fight continues

Instagram Threads is rolling out the ability for users to signal which sort of posts they wanted to see more or less of by swiping.

You can now customize your For You feed on Threads using swipes

The Japanese billionaire who commissioned SpaceX for a private mission around the moon on a Starship rocket has abruptly canceled the project, citing ongoing uncertainties around when the launch vehicle…

Japanese billionaire pulls plug on private ‘dearMoon’ lunar Starship mission

Malicious actors are abusing generative AI music tools to create homophobic, racist, and propagandic songs — and publishing guides instructing others how to do so. According to ActiveFence, a service…

People are using AI music generators to create hateful songs

As WWDC 2024 nears, all sorts of rumors and leaks have emerged about what iOS 18 and its AI-powered apps and features have in store.

What to expect from Apple’s AI-powered iOS 18 at WWDC

Dallas is the second city that Cruise is easing its way back into after pulling its entire U.S. fleet late last year.

GM’s Cruise is testing robotaxis in Dallas again

Featured Article

After raising $100M, AI fintech LoanSnap is being sued, fined, evicted

The company has been sued by at least seven creditors, including Wells Fargo.

15 hours ago
After raising $100M, AI fintech LoanSnap is being sued, fined, evicted

Featured Article

Sonos Ace review: A high-priced contender

The Ace are a contender in a crowded market, but they’re still in search of that magic bullet to truly let them stand out from the pack.

15 hours ago
Sonos Ace review: A high-priced contender

The change would see Instagram becoming more like the free version of YouTube, which requires users to view ads before and in the middle of watching videos.

Instagram confirms test of ‘unskippable’ ads

Commerce platform Shopify has acquired Checkout Blocks, allowing Shopify Plus merchants to make no-code customizations in their checkout to enhance customer experience and potentially boost sales.  Checkout Blocks, which debuted…

Shopify acquires Checkout Blocks, a checkout customization app

After the Digital Markets Act (DMA) forced Apple to allow third-party app stores for iOS in Europe, several developers have launched alternative stores, like the AltStore and MacPaw’s Setapp (currently…

Aptoide launches its alternative iOS game store in the EU

Time is relentless and, right now, it’s no friend to procrastination-prone early-stage startup founders. The application window for Startup Battlefield 200 (SB 200) at TechCrunch Disrupt 2024 slams shut in…

One week left: Apply to TC Disrupt Startup Battlefield 200

Cloudera, the once high-flying Hadoop startup, raised $1 billion and went public in 2018 before being acquired by private equity for $5.3 billion in 2021. Today, the company announced that…

Cloudera acquires Verta to bring some AI chops to its data platform

The global spend management sector is experiencing a tailwind of sorts. North America is arguably the biggest market in this space, but spend management companies have seen demand rise across…

Spend management startup SiFi raises $10M to grow further in Saudi Arabia

Neural Concept lets designers model how components will perform before they can be manufactured.

Swiss startup Neural Concept raises $27M to cut EV design time to 18 months

The StrictlyVC roadtrip continues! Coming off of sold-out events in London, Los Angeles, and San Francisco, we’re heading to Washington, D.C. for a cozy-vc-packed, evening at the Woolly Mammoth Theatre…

Don’t miss StrictlyVC in DC next week

X will now allow users to post consensually produced NSFW content as long as it is prominently labeled as such.

X tweaks rules to formally allow adult content

Ashby consolidates existing talent acquisition tools and leans heavily on AI to automate the more repetitive steps in the recruitment pipeline.

Ashby injects recruiting with a dose of AI

Spotify has announced it’s hiking subscriptions for customers in the U.S., the second such price increase in the space of a year. The music-streaming giant reports that premium pricing will…

Spotify to increase premium pricing in the US to $11.99 per month

Monzo has announced its 2024 financial results, revealing its first full-year pre-tax profit. The company also confirmed that it’s in the early stages of expanding into the broader European market…

UK neobank Monzo reports first full (pre-tax) profit, prepares for EU expansion with Dublin hub

Featured Article

Inside Apple’s efforts to build a better recycling robot

Last week, TechCrunch paid a visit to Apple’s Austin, Texas, manufacturing facilities. Since 2013, the company has built its Mac Pro desktop about 20 minutes north of downtown. The 400,000-square-foot facility sits in a maze of industry parks, a quick trip south from the company’s in-progress corporate campus. In recent years, the capital city has…

1 day ago
Inside Apple’s efforts to build a better recycling robot

Early attempts at making dedicated hardware to house artificial intelligence smarts have been criticized as, well, a bit rubbish. But here’s an AI gadget-in-the-making that’s all about rubbish, literally: Finnish…

Binit is bringing AI to trash

Temasek has previously invested in Lenskart, and this new funding follows a $500 million investment by the Abu Dhabi Investment Authority last year.

Temasek, Fidelity buy $200M stake in Lenskart at $5B valuation

Less than one year after its iOS launch, French startup ten ten has gone viral with a walkie talkie app that allows teens to send voice messages to their close…

French startup ten ten reinvents the walkie-talkie