← UPPER Resources · Talent Across Industries

How Does AI Sourcing Actually Work for IT Staffing — Without Scraping?

2025-09-22 · 8 min read

Alex Mercer
Alex Mercer
Chief Technology Officer
Compliant AI sourcing for IT staffing works through bring-your-own-credential access to authorized accounts and licensed data channels, not unauthorized scraping. This matters because the hiQ Labs v. LinkedIn case ended in a $500,000 judgment and injunction against hiQ on tort and terms-of-service grounds even after a partial CFAA win, and GDPR/CCPA both treat candidate profile data as personal data requiring a lawful basis — meaning scraping remains civilly and contractually risky even where it isn't criminally prosecutable.

Speed is the whole game in technical staffing, and it's tempting to assume the fastest way to build a candidate pool is the most aggressive one: scrape every public profile you can find, run it through a filter, start messaging. The legal record says this is a bad trade — not because scraping never works, but because the downside when it goes wrong is severe, and there's a compliant path that gets to the same passive candidates without the exposure.

What actually happened in the LinkedIn scraping case?

The most instructive precedent is hiQ Labs v. LinkedIn. After six years of litigation — including a Ninth Circuit ruling that initially sided with hiQ's right to access public profile data, remanded following the Supreme Court's Van Buren CFAA ruling — the case ended in December 2022 with a confidential settlement that included a $500,000 judgment against hiQ, a finding that hiQ was liable under California common-law torts of trespass to chattels and misappropriation, and an injunction barring hiQ from scraping LinkedIn going forward (Morgan Lewis legal analysis). The nuance matters: hiQ ultimately won on the criminal CFAA theory — scraping public data isn't automatically "hacking" — but still lost on tort grounds and under LinkedIn's Terms of Service. Scraping public profile data is not criminally prosecutable under CFAA, but it remains civilly and contractually risky, which is exactly the exposure a staffing agency doesn't want sitting on its balance sheet (Morgan Lewis; case summary).

Is scraping risk limited to recruiting specifically?

No — it's a broader pattern regulators are actively enforcing against. Clearview AI, a facial-recognition company built on scraped public images, was fined more than £7.5 million by the UK Information Commissioner's Office for unlawful data processing. The fine was later overturned on a jurisdictional technicality since Clearview served only non-UK/EU law enforcement clients, but the ICO explicitly stated the ruling does not remove its ability to act against international scraping companies processing UK residents' data (BBC News). It's a cautionary parallel, not a direct precedent, but it illustrates that regulatory exposure for scraping-based data businesses is a live, ongoing risk category — not a settled, closed chapter.

Where does GDPR/CCPA fit in?

Both frameworks treat professional profile data — names, employment history, skills, contact information — as personal data subject to lawful-basis, consent, and data-subject-rights requirements. Bulk scraping of candidate data without a clear lawful basis carries direct GDPR/CCPA exposure, independent of any platform Terms-of-Service breach. This is a well-established regulatory principle rather than a single citable statistic, and it applies regardless of whether the scraped data originated from LinkedIn, a job board, or any other public source.

So how does compliant AI sourcing actually reach passive candidates?

Through a bring-your-own-credential model: sourcing runs through the agency's own authorized accounts and licensed data channels, rather than scraping profiles without authorization. This avoids the exact terms-of-service and data-protection exposure that led to the hiQ judgment and remains structurally different from Clearview's unlawful-processing exposure, while still reaching the roughly 75% of the professional workforce that is passive and not actively job-searching (LinkedIn Talent Solutions). The mechanics look similar from the outside — continuous, multi-channel outreach to people who aren't applying to jobs — but the underlying data provenance is fundamentally different, and that difference is what keeps the agency's practices defensible if a candidate, platform, or regulator ever asks how a contact was sourced.

Does compliance mean sacrificing reach or speed?

Not structurally. The bottleneck in technical hiring isn't the existence of passive candidates — it's identifying and engaging them fast enough, given that top technical talent typically stays actively searchable for only about 10 days once in motion. A bring-your-own-credential sourcing engine that runs continuously across authorized channels can move at the same speed as an aggressive scraping-based tool, without carrying the tail risk of a multi-year lawsuit or a six-figure settlement sitting on the other side of a fast placement.

What should a technical staffing team actually check before adopting a sourcing tool?

Three questions cut through most vendor claims quickly: does the tool operate through authorized, licensed accounts rather than unauthorized scraping; does it provide a clear audit trail for how a given contact was sourced; and does its data-handling practice hold up under GDPR/CCPA lawful-basis requirements if a regulator or candidate ever asks. A vendor unable to answer any of the three plainly is very likely relying on scraped data, whatever marketing language is used to describe the underlying mechanism. Given that the hiQ Labs v. LinkedIn settlement took six years and ended in a $500,000 judgment despite a partial win on the criminal theory, the downside of getting this wrong isn't hypothetical — it's a documented, multi-year legal and reputational cost (Morgan Lewis).

UPPER's POV: Speed and compliance are not a trade-off in technical sourcing — they're the same design decision. UPPER works exclusively through authorized accounts and licensed channels rather than scraping, which keeps an agency's data practices defensible under GDPR/CCPA and platform terms while still reaching the passive majority of technical candidates that never show up in an applicant pool.

Key data points

References

  1. Morgan Lewis — LinkedIn v. hiQ legal analysis (scraping settlement, tort liability)
  2. Wikipedia — hiQ Labs v. LinkedIn case summary
  3. BBC News — Clearview AI UK ICO fine
  4. LinkedIn Talent Solutions — passive candidate research
  5. SHRM 2025 Benchmarking Report (technical candidate market window)

Read the interactive version: How Does AI Sourcing Actually Work for IT Staffing — Without Scraping?