sagentics.ai
sagentics.ai

Custom AI product and SaaS development

Profile enrichment platform for sales and recruiting: what it does and why POPIA changes the build

By SagenticsPublished

A profile enrichment platform takes a bare name and company and appends verified contact, firmographic, or candidate data so a sales or recruiting team can act on a lead or applicant without manual research. Sales enrichment and recruiting enrichment are not the same product, because B2B contact databases are built to index buyers, not the workforce, and in South Africa neither one works without a documented POPIA lawful basis for the enrichment itself. If you're deciding whether to buy a vendor API or build your own enrichment layer, that compliance gap is the first thing to solve, not the last.

What a profile enrichment platform actually does

The fill-the-gap model: from name and company to CRM-ready profile

Enrichment starts with a partial record, usually a name, an email, or a company domain, and fills in the rest: job title, seniority, direct phone or email, company size, industry, tech stack, or for recruiting, work history, skills, and current employer. The output gets pushed straight into a CRM, ATS, or outreach tool as a complete, ready-to-contact record. That's the whole value proposition: turn one data point into ten without a human doing the lookup.

Why single-source data caps out at 50-70% coverage

No single data provider covers the whole market. Contact databases are built from a mix of public web data, opt-in forms, partner feeds, and purchased lists, and each source has blind spots by geography, industry, or company size. Run a batch of 1,000 South African leads through any one vendor and you'll typically get a verified match on 50-70% of them. The rest come back empty or with stale data, which means your team is back to manual research for a third of the list.

Waterfall enrichment: cascading through providers to reach 85-95% coverage

The fix most serious platforms use is a waterfall: query provider A first, and for whatever comes back empty or low-confidence, query provider B, then C, until you've exhausted your source stack. Done well, this pushes coverage from the 50-70% single-source ceiling to 85-95%. It costs more per record because you're paying multiple vendors for the same query, but it's the difference between an enrichment tool that works on paper and one that actually reduces manual research hours.

Sales enrichment and recruiting enrichment are not the same product

Why B2B contact databases index buyers, not candidates

Sales enrichment vendors built their databases by crawling and buying data on people who buy things: decision-makers, procurement contacts, marketing and IT leads. That data is deep on job titles and buying intent and shallow on career history, skills, and availability. A recruiting team needs the opposite: full work history, tenure patterns, skill tags, and signals on whether someone is open to moving. Those are structurally different datasets, sourced differently, and a platform optimised for one is weak on the other.

The buying mistake: recruiters sold a sales tool with a recruiting label

A lot of recruiting teams end up on a sales enrichment tool because the vendor added a "recruiting" tab to an existing sales product. The result is thin coverage on mid-level and technical candidates, almost nothing on passive candidates outside LinkedIn's crawl, and contact data that skews toward whoever appears in a sales-facing database. If you're evaluating a tool for talent sourcing, ask directly what percentage of their database is sourced from employment records and professional networks versus marketing and sales lead lists. If they can't answer, it's a sales tool wearing a recruiting badge.

What a candidate enrichment API needs that a sales tool doesn't

A genuine candidate enrichment API needs employment history depth, skill and certification tagging, education verification where relevant, and ideally some signal on career trajectory or flight risk. It also needs to handle consent differently, because appending personal information to a candidate profile for recruitment purposes is a different lawful basis question than appending a business email to a prospect record. This is where our own work overlaps with the AI candidate assessment in executive search we've built for clients: enrichment is only useful if the downstream assessment logic can actually use the extra fields, not just display them.

Getting Your POPIA Ducks in a Row - GoodX Software

Compliance and data provenance are now the competitive moat

What happened to Proxycurl and why scraping-based providers are a liability

Proxycurl, a widely used LinkedIn scraping API for enrichment platforms, was shut down in 2024 after LinkedIn's parent company pursued legal action over unauthorised scraping and terms-of-service violations. Platforms that had built enrichment pipelines on top of it lost their data source overnight, and some had to rebuild core features from scratch. That's not a compliance footnote. It's a business continuity risk: if your enrichment vendor sources data by scraping a platform that explicitly prohibits it, you inherit that legal exposure the moment their access gets cut off.

How to vet an enrichment vendor's sourcing method before you integrate it

Before you integrate any enrichment API, ask three questions in writing: where does the data come from, is it obtained under a licence or partnership agreement, and what happens to your pipeline if that source disappears. A vendor that can point to opt-in data collection, licensed partner feeds, or first-party consent flows is a materially different risk profile than one that's vague about sourcing or leans on "publicly available" as the entire answer. Get this in the contract, not just the sales call.

POPIA and enrichment: the lawful basis question nobody in this space is answering

Every enrichment vendor writing about compliance right now is writing for GDPR and CCPA. Almost none of them mention POPIA, and that's the gap South African sales and recruiting teams are walking into unaware. Under POPIA, appending personal information to a record, whether it's a prospect's mobile number or a candidate's employment history, is processing, and processing needs a lawful basis: consent, contractual necessity, legitimate interest properly weighed, or a legal obligation. Running a US-built enrichment API against a list of South African leads or candidates without documenting which basis applies is processing personal information without a defensible position if a data subject complains or the Information Regulator asks. This is the exact gap the profile intelligence platform we built was designed to close: POPIA consent and cross-border transfer restrictions are handled as an architecture decision at the pipeline level, not a legal disclaimer added after the product ships.

Data decay and how often you actually need to re-enrich

Why contact and candidate data decays 2-3 times a year

People change jobs, phone numbers, and email addresses constantly. Industry data consistently shows B2B contact records decay at roughly 25-30% a year, which in practice means a meaningful share of your enriched database goes stale two to three times annually. Candidate data decays on a similar curve, driven by job changes, skill updates, and availability shifts. A one-time enrichment run has a shelf life measured in months, not years.

Continuous enrichment vs 90-day batch refresh

Two models handle this: continuous enrichment, where records are refreshed in near real time as triggers fire (a job change detected, a company funding event), and batch refresh, where the whole database gets re-run against enrichment providers on a fixed cycle, typically every 90 days. Continuous enrichment costs more to build and run but keeps outreach data current. Batch refresh is cheaper and simpler to operate but means you're working with data that's, on average, 45 days stale at any given point in the cycle.

What stale data costs a sales or recruiting pipeline

Stale data doesn't just waste time, it actively damages outreach. A sales rep contacting someone who left the company six months ago burns the call and, worse, can flag your domain for poor list hygiene. A recruiter reaching out about a role the candidate accepted elsewhere last quarter looks careless. Teams that don't budget for re-enrichment usually discover the cost in declining reply rates and rising unsubscribe or complaint rates long before they trace it back to data age.

SURVEY: Digitisation and secure data destruction key elements of POPIA  compliance | ITWeb

Why single-source matching has an accuracy ceiling

The 75-85% structural cap on entity resolution

Matching a partial record to the right person, entity resolution, hits a structural accuracy ceiling around 75-85% when you rely on a single data source, because names and companies aren't unique identifiers and providers use different matching keys (email domain, phone hash, LinkedIn URL). Waterfall enrichment across multiple providers pushes this higher by cross-validating matches, but no provider or combination gets to 100%, and any platform claiming it should be treated with suspicion.

Merge errors and what they cost in wrong-person outreach

The cost of matching errors isn't abstract. A wrong-person merge means your CRM shows the wrong job title, the wrong company, or in the worst case, merges two different people's history into one record. For sales, that's an embarrassing email to the wrong contact. For recruiting, it can mean surfacing the wrong person's employment history in a background reference check, which is a POPIA problem as well as a credibility one. Build in confidence scores on every enriched field and route low-confidence matches to manual review rather than auto-populating them.

Building an enrichment platform in South Africa: what it actually takes

API-first architecture for platform builders vs Chrome-extension tools

If you're building a product that other tools will plug into, an API-first architecture is non-negotiable. Chrome extension enrichment tools (the kind that scrape a LinkedIn profile page you're viewing) are fine for individual users doing manual lookups, but they don't scale, don't batch process, and can't be embedded into a CRM or ATS workflow. A platform aimed at other builders needs a proper API: rate limits, webhooks for async enrichment jobs, confidence scores on every field, and clear documentation on data provenance per field, not just per record.

Where POPIA changes the build compared to a GDPR-first vendor

A GDPR-first vendor architecture assumes EU-based consent management and EU-to-anywhere transfer rules. POPIA has its own requirements around cross-border transfer of personal information (Section 72) and a stricter default toward consent for direct marketing use of enriched contact data. That means the build needs a jurisdiction flag on every record, different retention and consent-logging logic for South African data subjects, and an audit trail that can produce, on request, exactly which lawful basis applied to a given enrichment action. Retrofitting this after the fact is expensive. Building it in from day one is a schema decision, and it's covered in more depth in what actually needs to be true for POPIA compliance for automated systems handling personal information generally.

Build vs buy: when a custom enrichment layer is worth it

Buying a vendor API makes sense if your compliance exposure is low, your volume is modest, and you can get contractual guarantees on sourcing and POPIA handling. Building your own enrichment layer makes sense when you're processing candidate or prospect data at scale, when vendor compliance answers are vague, or when enrichment is core to your product rather than a feature bolt-on. We've laid out the honest build vs buy answer in more detail, but the short version for enrichment specifically: if POPIA lawful basis isn't something your vendor can document in writing, you're already carrying build-level risk without build-level control.

How Sagentics structures enrichment for POPIA compliance

We treat consent and lawful basis as a field in the schema, not a policy document sitting outside the system. Every enriched record carries a provenance tag (which source, under what basis, when), cross-border transfer flags are set at ingestion, and retention windows are enforced automatically rather than left to manual cleanup. That's the architecture behind the profile intelligence platform we built, and it's the same approach we bring to any enrichment work, whether it's feeding a sales CRM or a recruiting pipeline.

Where this fits alongside AI candidate assessment work

Enrichment is rarely the end goal, it's the input to a decision: who to call, who to shortlist, who to assess further. We've built this pipeline end to end for executive search, where enriched candidate profiles feed directly into structured AI assessment, documented in our executive candidate assessment case study. If you're scoping a similar build, it's worth reading through scoping an AI SaaS build before you pay for it first, and our broader work on custom AI development in South Africa covers how we approach these systems end to end.

Common questions

What is profile or customer enrichment and how does it differ from basic contact data? Basic contact data is whatever you collected directly, usually a name and email. Enrichment appends everything else, job title, company, phone, firmographics, or candidate history, by matching that partial record against external data providers. The difference is completeness: enrichment turns a lead form submission into a fully qualified, actionable record without manual research.

What are the best candidate or recruiting enrichment APIs for platform builders? There's no single best option, since coverage varies sharply by industry and seniority level. Evaluate on three criteria: what percentage of their data is sourced from employment records rather than sales lead lists, how they handle POPIA-relevant consent for South African candidates, and whether they support waterfall matching with confidence scoring rather than a single opaque match.

How often should candidate or contact records be re-enriched? Most records need refreshing two to three times a year, since contact and employment data decays roughly 25-30% annually. Batch refresh every 90 days is a reasonable default for most pipelines. High-value or high-velocity pipelines, like active executive search, benefit from continuous enrichment triggered by detected job or company changes instead.

Is scraped LinkedIn data compliant, and what happened to Proxycurl? Proxycurl, a major LinkedIn scraping API, was shut down in 2024 following legal action from LinkedIn's parent company over unauthorised scraping. Scraped data sits in a legal grey zone at best and violates platform terms at worst, and vendors relying on it carry both compliance risk and business continuity risk if their data source disappears without warning.

Which enrichment tool is right for recruiters vs sales teams? Sales enrichment tools are built on databases indexing buyers and decision-makers, deep on job titles and company data but shallow on career history. Recruiting needs a tool sourced from employment records with skill tagging and work history depth. Buying a relabelled sales tool for recruiting means weak coverage on technical and passive candidates.

What's the difference between an enrichment API and an enrichment platform or Chrome extension? A Chrome extension enriches one profile at a time as you browse, useful for individual manual lookups but not scalable. An API-first platform processes records in bulk, integrates into a CRM or ATS via webhooks, and returns structured, confidence-scored data suitable for automated pipelines rather than one-off checks.

How does POPIA affect using enrichment data on South African candidates or prospects? Appending personal information to a record counts as processing under POPIA, which needs a lawful basis, consent, legitimate interest properly assessed, or contractual necessity. Most enrichment vendors are built for GDPR and don't address this. South African teams using them are frequently processing personal information without a documented basis, which is a real exposure if challenged.

How much time does enrichment actually save a sales or recruiting team? Manual research on a single lead or candidate typically takes 5-15 minutes across LinkedIn, company sites, and search. Enrichment collapses that to seconds per record at scale. For a team processing hundreds of leads or candidates a week, that's realistically 10-20 hours a week returned to actual selling or sourcing work rather than lookup.

If you're weighing up whether to buy an enrichment vendor or build your own POPIA-aware layer, message us on WhatsApp and we'll talk through what actually fits your pipeline.

Common questions

What is profile or customer enrichment and how does it differ from basic contact data?

Basic contact data is whatever you collected directly, usually a name and email. Enrichment appends everything else: job title, company, phone, firmographics, or candidate history, by matching that partial record against external data providers. The difference is completeness: enrichment turns a lead form submission into a fully qualified, actionable record without manual research.

What are the best candidate or recruiting enrichment APIs for platform builders?

There's no single best option, since coverage varies sharply by industry and seniority level. Evaluate on three criteria: what percentage of their data is sourced from employment records rather than sales lead lists, how they handle POPIA-relevant consent for South African candidates, and whether they support waterfall matching with confidence scoring rather than a single opaque match.

How often should candidate or contact records be re-enriched?

Most records need refreshing two to three times a year, since contact and employment data decays roughly 25-30% annually. Batch refresh every 90 days is a reasonable default for most pipelines. High-value or high-velocity pipelines, like active executive search, benefit from continuous enrichment triggered by detected job or company changes instead.

Is scraped LinkedIn data compliant, and what happened to Proxycurl?

Proxycurl, a major LinkedIn scraping API, was shut down in 2024 following legal action from LinkedIn's parent company over unauthorised scraping. Scraped data sits in a legal grey zone at best and violates platform terms at worst, and vendors relying on it carry both compliance risk and business continuity risk if their data source disappears without warning.

Which enrichment tool is right for recruiters vs sales teams?

Sales enrichment tools are built on databases indexing buyers and decision-makers, deep on job titles and company data but shallow on career history. Recruiting needs a tool sourced from employment records with skill tagging and work history depth. Buying a relabelled sales tool for recruiting means weak coverage on technical and passive candidates.

What's the difference between an enrichment API and an enrichment platform or Chrome extension?

A Chrome extension enriches one profile at a time as you browse, useful for individual manual lookups but not scalable. An API-first platform processes records in bulk, integrates into a CRM or ATS via webhooks, and returns structured, confidence-scored data suitable for automated pipelines rather than one-off checks.

How does POPIA affect using enrichment data on South African candidates or prospects?

Appending personal information to a record counts as processing under POPIA, which needs a lawful basis: consent, legitimate interest properly assessed, or contractual necessity. Most enrichment vendors are built for GDPR and don't address this. South African teams using them are frequently processing personal information without a documented basis, which is a real exposure if challenged.

How much time does enrichment actually save a sales or recruiting team?

Manual research on a single lead or candidate typically takes 5-15 minutes across LinkedIn, company sites, and search. Enrichment collapses that to seconds per record at scale. For a team processing hundreds of leads or candidates a week, that's realistically 10-20 hours a week returned to actual selling or sourcing work rather than lookup.

About Sagentics

Sagentics is an AI systems studio based in South Africa. We design and build WhatsApp automation, n8n workflows, and custom AI products for local and international clients. We write from systems we have actually shipped.

Start a WhatsApp conversation with Sagentics

Related reading