The client is a job search platform, where people need to see the actual pay and the location of the jobs before they apply.
Every listing starts as a link on an employer's career page. This automation collects them from 300+ sites, drops the duplicates, reads each posting with AI, and files finished records into the client's database.
Their team once did all of it by hand. Within the month I started the project their output had already multiplied, and by the second month it had more than quadrupled.
The team's target is at least 500 finished links a day. By hand that is 3 to 4 hours a batch and sometimes 8, spent switching tabs and copy pasting. The same 500 now run themselves, and the person's whole job is checking the output at the end: about 30 minutes.
Open 300+ career sites one by one, sometimes with a VPN, and copy paste every link into a sheet.
Search the database for every collected link. Most turn out to be repeats.
Open every surviving link and copy 10+ fields per job, one field at a time. The description and qualifications also have to be rewritten to a set format, every single job.
Hunt the exact address in a masterlist, job by job. The most tedious step of all.
One click from the sheet menu collects every job link from all 300+ career sites, allowed-state jobs only.
The sheet flags repeats by colour, then the Main Filter lets only truly new links through.
AI opens each job page, fills in every field and writes the description and qualifications in the required format, one record every 30 seconds.
The exact address fills itself from a 95,000-row reference database, then clean records land in the database. Slack pings me if anything fails.
Four rules decided what got automated first, how it holds up across 300+ sites, and how I know when something breaks.
Address hunting and field-by-field scraping were the most tedious, so they went first and paid back soonest. Link collecting stayed manual for three months on purpose: it was the biggest build, and the client kept working the whole time. Nothing waited for a finished system. Each piece went live as it was done, so the team was saving hours from the first month. Opening every career site by hand, sometimes with a VPN, is now one click.
The 300+ career sites run on about 40 hiring platforms, and every one is its own small puzzle: pagination, region locks, and each company's own rules about which US states are allowed. Each platform gets its own handler, so a site that changes or breaks never takes the pipeline down, and a new site gets added without touching the rest.
A silent failure is worse than a stopped one, because bad records are trusted. Every row carries a tracking ID, the AI's output is checked before anything saves, and Slack pings me the moment a step fails. Nothing rots in the database unnoticed.
I tried the same build in Make and in Zapier first. Both cost far more at this volume and neither could handle the branching, retries and per-platform logic this needed. n8n, self-hosted, was the only one flexible enough. The control panel is a Google Sheet on purpose: the team already worked there, so nobody had to learn a new tool.
A system only one person can run is a liability. I trained the client's staff to operate and troubleshoot it themselves, so the day to day is theirs. I check the workflows about once a week, step in when a site change is beyond them, and answer questions any time.
An automation that quietly overspends is only half built. I checked the scraping bill against what the pipeline actually needed and found every request was being charged at the most expensive tier. Fixing it cut scraping costs by 42%, with no change to coverage or results. In July the whole pipeline ran on about $3 per 1,000 finished records: roughly $2.20 of AI reading and $0.80 of page fetching, for 19,796 records.
No build this size runs clean from the start. Three things went wrong, and each one changed how the system works.
A career site would redesign a page or add bot protection, and its handler would quietly return nothing. That is why every step now reports failures to Slack instead of just moving on, and why handlers are separate: one site going dark never stops the other 299.
Early on, some records were returned by the client for falling short of their standard. I tightened the prompt and added the check step that reads the AI's output before anything saves. Returns dropped and have stayed low since.
The scraping service was charging every request at its most expensive tier, whether the page needed it or not. I only found it by checking the bill against what the pipeline actually did. Fixing it cut scraping cost by 42%, and I have watched the run cost since.
The seven-step journey from a company career page to a clean, address-verified record in Airtable.
One click gathers every job link from 300+ career sites, allowed-state jobs only.
Colour-codes duplicate links, gives each row a unique tracking ID, and sends the clean batch to the next step.
Every link is checked against the full database. Only truly new ones join the queue.
Tries three ways to load each page in order, from free to paid, escalating automatically only when needed.
AI pulls out the title, pay, location, and requirements. Low-pay postings are skipped before AI even runs.
Partial addresses are completed from the reference database.
The AI's output is double-checked, then the clean record saves to Airtable.
A 90-second run: duplicate flagging in the sheet, the filters and scraper in n8n, and records landing in Airtable.
The guardrails built into the pipeline.
Every link gets a permanent tracking ID, so updates always land on the right record.
Repeats are caught in the sheet first, then against the full masterlist. No duplicate reaches the queue.
The 300+ career sites run on about 40 hiring platforms: Taleo, ADP, UKG, Dayforce, iCIMS, Greenhouse, Workable and more. Each platform is fetched its own way.
The client can bump any company to the front of the next run and watch progress live in the sheet.
Each page is tried the free way first. Paid tools step in only when a page blocks them.
AI reads the posting and fills the fields. A second step checks the result before it saves.
Partial addresses are completed automatically from a 95,000-row reference database.
Obviously low-pay postings are skipped before the AI runs, so it never spends on a dead end.
A new run exits if the previous one is still going, so nothing is processed twice.
Screenshots of all workflows, showing every step, branch, and error path in the canvas.
From the sheet menu, it walks all 300+ career sites and collects every job link in one run, even pages listing 999+ openings.
The 300+ sites run on about 40 hiring platforms, each fetched its own way. Sites the team used to reach with a VPN now go through a paid scraping service instead.
Only allowed-state jobs are kept, new links land in the sheet, and each company's visit date is stamped automatically.
Four colours mark four kinds of repeat, from already-processed links to same-day resubmissions, so every row's state is visible at a glance.
Every new row gets a permanent tracking number the rest of the system uses to follow it.
One click sends all pending links to the next step and re-sorts the sheet, oldest first.
Reads the submitted links, strips tracking junk from each URL, and de-duplicates the batch.
Each link is checked against the full job database. If it was submitted before, the sheet row is updated to show that. If it is new, it moves forward.
New links join the processing queue and the row is marked Transferred. Any error marks the row and pings Slack.
Checks that no previous run is still going, then rewrites platform links into a directly fetchable format.
A standard request first, a rendering tool if the page needs JavaScript, and a paid tool only as a last resort.
Obviously low-pay postings are skipped at no cost. Everything else is read by AI, field by field.
Extracted salaries are sanity-checked. Records below the pay threshold are marked skipped, not qualified.
Each record saves as Qualified, Salary Issue, or Error. Incomplete addresses are flagged for the Find Address workflow.
Every skipped record sends a Slack message with the reason. Qualified records save silently.
Finds all records in Airtable that the scraper flagged as having an incomplete address and processes them one by one.
Breaks the partial address into pieces and searches the 95,000-row database for companies that match by name, returning the top candidates to score.
Candidates are ranked by closeness of match, with a backup lookup if nothing hits. Worst case, the original partial address is kept.
Reads the master address sheet and prepares every row for upload.
Uploads in batches. Existing rows are skipped, so re-running never creates duplicates.
Reports how many rows were added and how many skipped. Run it whenever the masterlist changes.
Reads the companies list in the sheet and fetches the matching records from the shared database, so the two can be compared in one pass.
Compares the two and updates only the records that actually changed, keeping the sync fast and avoiding unnecessary writes.
A small companion to the Link Collector: the client's team always sees fresh last-updated dates for every company, with no manual logging.
Every tool in the pipeline, and exactly where AI fits in.
In the pipeline: an AI model (DeepSeek V4 Flash) reads job pages and fills in the fields. DeepSeek was chosen after testing several models head to head: it gave the most accurate output for the lowest cost per job. Every AI output passes rule checks before it is saved.
Two databases, on purpose: Airtable holds the job records, and Supabase holds only the 95,000-row address reference. Address lookups run on every single job, so they need a database that answers fast and does not run into row limits. Airtable stays where the team already works.
In the build: I used Claude (Anthropic) as an assistant for drafting code and copy. The architecture, decisions, testing, and client work are mine, and I review everything the AI touches before it ships.
Completed job links are the team's core monthly output number. The automation went live at the start of May. April was a genuinely slow month for the team, and the manual figures are shown exactly as recorded.
Finished records are fed into the client's own database, and anything that falls short of their standard is returned to us. Returns are rare, and none of them are counted above. There is no better accuracy check than this one: the client pays on accepted records, so every number here is output they paid for. The pipeline is still running today. The chart stops at July only because this case study does.
This one was niche. My client thought so too when we started. Tell me the manual work eating your time and we will find a way and the right tools to make it run itself.
Discovery is a free scoped audit, back within 24 to 48 hours, no calls needed. Planning maps every field and step as an SOP with a fixed quote attached, and nothing gets built until you approve it. Then Build, then Handover. Hourly support starts at $15/hr.
The workflows and the documentation are yours, written in plain English. Some clients want it in their own n8n and run it themselves. Others would rather not carry the maintenance and leave it with me. Both are fine, and we settle which in planning.
On a retainer I am there when something goes wrong. Off one, you are not stranded: the documentation covers it, and I train your staff to handle the day to day themselves, the same way I did here. Questions are always welcome either way.
You do not need to arrive with a plan, a tool list, or any technical detail. Describe the manual work eating your time and I take it from there. Anything I need from you, accounts, access, a reference file, a person to check output, is listed in discovery and planning before a quote, so there are no surprises later.
Open to automation projects, contract work, and full-time positions.