← Portfolio  /  Case Study
Case Study · Job Link Automation

JobLinkOS · From raw job link to qualified record, automated end-to-end

The client is a job search platform, where people need to see the actual pay and the location of the jobs before they apply.

Every listing starts as a link on an employer's career page. This automation collects them from 300+ sites, drops the duplicates, reads each posting with AI, and files finished records into the client's database.

Their team once did all of it by hand. Within the month I started the project their output had already multiplied, and by the second month it had more than quadrupled.

Stack · n8n · Google Apps Script · Airtable · Supabase · AI extraction API · paid scraping API Role · Sole automation engineer: architecture, build, maintenance, and training the client's team to run it Note · JobLinkOS is my own system. The business running it is a client, and no client data is shown Sources · The 300+ career sites belong to the client's own partner employers. When a company stops working with them, its site comes off the list
01 · The Client's Problem

Four manual chores became one automated pipeline

The team's target is at least 500 finished links a day. By hand that is 3 to 4 hours a batch and sometimes 8, spent switching tabs and copy pasting. The same 500 now run themselves, and the person's whole job is checking the output at the end: about 30 minutes.

✕ Before · All by hand
The same four chores, every single batch

1 · Link collecting

Open 300+ career sites one by one, sometimes with a VPN, and copy paste every link into a sheet.

↓

2 · Links checking

Search the database for every collected link. Most turn out to be repeats.

↓

3 · Data scraping

Open every surviving link and copy 10+ fields per job, one field at a time. The description and qualifications also have to be rewritten to a set format, every single job.

↓

4 · Find address

Hunt the exact address in a masterlist, job by job. The most tedious step of all.

3 to 4 hours per batch, sometimes 8, every batch
✓ After · JobLinkOS
The same four chores, run from one sheet

1 · Link Collector New

One click from the sheet menu collects every job link from all 300+ career sites, allowed-state jobs only.

↓

2 · Duplicate filters

The sheet flags repeats by colour, then the Main Filter lets only truly new links through.

↓

3 · Data Scraper

AI opens each job page, fills in every field and writes the description and qualifications in the required format, one record every 30 seconds.

↓

4 · Find Address + Sync

The exact address fills itself from a 95,000-row reference database, then clean records land in the database. Slack pings me if anything fails.

About 30 minutes, all of it checking the output
02 · My Approach

Built in order of pain, and built to keep running

Four rules decided what got automated first, how it holds up across 300+ sites, and how I know when something breaks.

Worst chore first, link collecting last

Address hunting and field-by-field scraping were the most tedious, so they went first and paid back soonest. Link collecting stayed manual for three months on purpose: it was the biggest build, and the client kept working the whole time. Nothing waited for a finished system. Each piece went live as it was done, so the team was saving hours from the first month. Opening every career site by hand, sometimes with a VPN, is now one click.

Was the last manual stepNow 1 click

Every site gets its own handler

The 300+ career sites run on about 40 hiring platforms, and every one is its own small puzzle: pagination, region locks, and each company's own rules about which US states are allowed. Each platform gets its own handler, so a site that changes or breaks never takes the pipeline down, and a new site gets added without touching the rest.

TaleoADPUKGDayforceiCIMSGreenhouse+ more

It fails loudly, never quietly

A silent failure is worse than a stopped one, because bad records are trusted. Every row carries a tracking ID, the AI's output is checked before anything saves, and Slack pings me the moment a step fails. Nothing rots in the database unnoticed.

Slack alertsRow tracking IDsOutput checked before save

n8n because the others could not do it

I tried the same build in Make and in Zapier first. Both cost far more at this volume and neither could handle the branching, retries and per-platform logic this needed. n8n, self-hosted, was the only one flexible enough. The control panel is a Google Sheet on purpose: the team already worked there, so nobody had to learn a new tool.

Make: too costly at volumeZapier: too rigidn8n, self-hosted

Handed over, not hoarded

A system only one person can run is a liability. I trained the client's staff to operate and troubleshoot it themselves, so the day to day is theirs. I check the workflows about once a week, step in when a site change is beyond them, and answer questions any time.

Client's staff run itWeekly checkOn call for questions

I watch what it costs to run

An automation that quietly overspends is only half built. I checked the scraping bill against what the pipeline actually needed and found every request was being charged at the most expensive tier. Fixing it cut scraping costs by 42%, with no change to coverage or results. In July the whole pipeline ran on about $3 per 1,000 finished records: roughly $2.20 of AI reading and $0.80 of page fetching, for 19,796 records.

~$3 per 1,000 records42% lower scraping costSame coverage

What broke, and what I changed

No build this size runs clean from the start. Three things went wrong, and each one changed how the system works.

Sites changed without warning

A career site would redesign a page or add bot protection, and its handler would quietly return nothing. That is why every step now reports failures to Slack instead of just moving on, and why handlers are separate: one site going dark never stops the other 299.

Records came back rejected

Early on, some records were returned by the client for falling short of their standard. I tightened the prompt and added the check step that reads the AI's output before anything saves. Returns dropped and have stayed low since.

The bill was quietly wrong

The scraping service was charging every request at its most expensive tier, whether the page needed it or not. I only found it by checking the bill against what the pipeline actually did. Fixing it cut scraping cost by 42%, and I have watched the run cost since.

03 · The Pipeline

From link entry to qualified record

The seven-step journey from a company career page to a clean, address-verified record in Airtable.

1 · Collect the links

One click gathers every job link from 300+ career sites, allowed-state jobs only.

2 · Sheet entry and cleanup

Colour-codes duplicate links, gives each row a unique tracking ID, and sends the clean batch to the next step.

3 · Already-seen check

Every link is checked against the full database. Only truly new ones join the queue.

4 · Open the job page

Tries three ways to load each page in order, from free to paid, escalating automatically only when needed.

5 · AI reads the posting

AI pulls out the title, pay, location, and requirements. Low-pay postings are skipped before AI even runs.

6 · Fill in the address

Partial addresses are completed from the reference database.

7 · Check and save

The AI's output is double-checked, then the clean record saves to Airtable.

04 · Live Demo

Watch the pipeline run

A 90-second run: duplicate flagging in the sheet, the filters and scraper in n8n, and records landing in Airtable.

Heads up: this video shows an earlier version of the system, before the Link Collector was added. An updated walkthrough is coming soon.
05 · The Guardrails

Built so nothing slips through

The guardrails built into the pipeline.

🆔

Unique ID per Row

Every link gets a permanent tracking ID, so updates always land on the right record.

🔍

Two-Layer Duplicate Check

Repeats are caught in the sheet first, then against the full masterlist. No duplicate reaches the queue.

🔧

40 Platforms, 300+ Sites

The 300+ career sites run on about 40 hiring platforms: Taleo, ADP, UKG, Dayforce, iCIMS, Greenhouse, Workable and more. Each platform is fetched its own way.

⚡

Priority Queue and Live Progress

The client can bump any company to the front of the next run and watch progress live in the sheet.

📡

Three-Way Page Loading

Each page is tried the free way first. Paid tools step in only when a page blocks them.

🤖

AI Reading with Output Check

AI reads the posting and fills the fields. A second step checks the result before it saves.

📍

Address Completion

Partial addresses are completed automatically from a 95,000-row reference database.

💰

Cost Control Built In

Obviously low-pay postings are skipped before the AI runs, so it never spends on a dead end.

🛡️

No-Overlap Protection

A new run exits if the previous one is still going, so nothing is processed twice.

06 · The Tools

Every workflow, fully wired

Screenshots of all workflows, showing every step, branch, and error path in the canvas.

Link Collector: Every Career Site, One Click

n8n · Menu-triggered · The last manual step, automated
Link Collector: n8n workflow canvas
1
One Click, Every Site

From the sheet menu, it walks all 300+ career sites and collects every job link in one run, even pages listing 999+ openings.

2
A Strategy per Platform

The 300+ sites run on about 40 hiring platforms, each fetched its own way. Sites the team used to reach with a VPN now go through a paid scraping service instead.

3
Filter, Append, and Stamp

Only allowed-state jobs are kept, new links land in the sheet, and each company's visit date is stamped automatically.

Initial Filter: Google Sheet Automation Menu

Apps Script · Menu-triggered
Google Sheet Automation menu
1
Four-Colour Duplicate Flagging

Four colours mark four kinds of repeat, from already-processed links to same-day resubmissions, so every row's state is visible at a glance.

2
Tracking ID Assignment

Every new row gets a permanent tracking number the rest of the system uses to follow it.

3
Send to Automation and Sort

One click sends all pending links to the next step and re-sorts the sheet, oldest first.

Main Filter: Already-Seen Check and Queue Push

Triggered automatically · Already-seen check
Job Link OS Main Filter: n8n workflow canvas
1
Read and Clean Up

Reads the submitted links, strips tracking junk from each URL, and de-duplicates the batch.

2
Already-Seen Check

Each link is checked against the full job database. If it was submitted before, the sheet row is updated to show that. If it is new, it moves forward.

3
Push and Confirm

New links join the processing queue and the row is marked Transferred. Any error marks the row and pings Slack.

Data Scraper: Open, Read, and Save

Runs on a schedule · AI extraction to Airtable
Data Scraper: n8n workflow canvas
1
No Double-Runs and Link Cleanup

Checks that no previous run is still going, then rewrites platform links into a directly fetchable format.

2
Three-Way Page Loading

A standard request first, a rendering tool if the page needs JavaScript, and a paid tool only as a last resort.

3
Salary Pre-check and AI Reading

Obviously low-pay postings are skipped at no cost. Everything else is read by AI, field by field.

4
Output Check

Extracted salaries are sanity-checked. Records below the pay threshold are marked skipped, not qualified.

5
Save to Airtable

Each record saves as Qualified, Salary Issue, or Error. Incomplete addresses are flagged for the Find Address workflow.

6
Slack Notifications

Every skipped record sends a Slack message with the reason. Qualified records save silently.

Find Address: Completing Partial Addresses

Runs on a schedule · Reference database lookup
Find Address workflow: n8n canvas
1
Pick Up Incomplete Records

Finds all records in Airtable that the scraper flagged as having an incomplete address and processes them one by one.

2
Search the Reference Database

Breaks the partial address into pieces and searches the 95,000-row database for companies that match by name, returning the top candidates to score.

3
Best-Match Scoring with Fallback

Candidates are ranked by closeness of match, with a backup lookup if nothing hits. Worst case, the original partial address is kept.

Sync Database: Keep the Address Reference Up to Date

Run manually as needed · 95,000 rows
Sync Database workflow: n8n canvas
1
Read the Address Sheet

Reads the master address sheet and prepares every row for upload.

2
Upload to the Reference Database

Uploads in batches. Existing rows are skipped, so re-running never creates duplicates.

3
Sync Summary

Reports how many rows were added and how many skipped. Run it whenever the masterlist changes.

Companies Sync: Keep the Companies List and Database in Step

n8n · Supporting workflow · New
Companies Sync: n8n workflow canvas
1
Read Both Sides

Reads the companies list in the sheet and fetches the matching records from the shared database, so the two can be compared in one pass.

2
Update Only What Changed

Compares the two and updates only the records that actually changed, keeping the sync fast and avoiding unnecessary writes.

3
Fresh Dates for the Team

A small companion to the Link Collector: the client's team always sees fresh last-updated dates for every company, with no manual logging.

07 · The Stack

The stack, and how AI was used

Every tool in the pipeline, and exactly where AI fits in.

AI Transparency

In the pipeline: an AI model (DeepSeek V4 Flash) reads job pages and fills in the fields. DeepSeek was chosen after testing several models head to head: it gave the most accurate output for the lowest cost per job. Every AI output passes rule checks before it is saved.

Two databases, on purpose: Airtable holds the job records, and Supabase holds only the 95,000-row address reference. Address lookups run on every single job, so they need a database that answers fast and does not run into row limits. Airtable stays where the team already works.

In the build: I used Claude (Anthropic) as an assistant for drafting code and copy. The architecture, decisions, testing, and client work are mine, and I review everything the AI touches before it ships.

n8n (self-hosted, Hostinger VPS) Google Apps Script Google Sheets Airtable · every record Supabase / PostgreSQL · address lookups AI extraction (DeepSeek V4 Flash) Jina AI Reader Paid scraping API (rendered pages) Slack GitHub (workflow JSON exports)
08 · The Outcome

Job links completed per month

Completed job links are the team's core monthly output number. The automation went live at the start of May. April was a genuinely slow month for the team, and the manual figures are shown exactly as recorded.

Manual Automated
20k 15k 10k 5k January · 3,565 job links · manual February · 3,411 job links · manual March · 3,826 job links · manual April · 2,870 job links · manual May · 12,599 job links · automated June · 14,822 job links · automated July · 19,786 job links · automated 3,565 3,411 3,826 2,870 12,599 14,822 19,786 5.8x the manual average Jan Feb Mar Apr May Jun Jul
June ran with the client mostly away from the desk. The pipeline kept collecting, checking, and filling records on its own.
Every number here is an accepted record

Finished records are fed into the client's own database, and anything that falls short of their standard is returned to us. Returns are rare, and none of them are counted above. There is no better accuracy check than this one: the client pays on accepted records, so every number here is output they paid for. The pipeline is still running today. The chart stops at July only because this case study does.

09 · Working Together

What a build like this looks like for you

This one was niche. My client thought so too when we started. Tell me the manual work eating your time and we will find a way and the right tools to make it run itself.

Four steps, fully async

Discovery is a free scoped audit, back within 24 to 48 hours, no calls needed. Planning maps every field and step as an SOP with a fixed quote attached, and nothing gets built until you approve it. Then Build, then Handover. Hourly support starts at $15/hr.

DiscoveryPlanningBuildHandover

You own what gets built

The workflows and the documentation are yours, written in plain English. Some clients want it in their own n8n and run it themselves. Others would rather not carry the maintenance and leave it with me. Both are fine, and we settle which in planning.

Full documentationYour accountsYour choice of tools

What happens when it breaks

On a retainer I am there when something goes wrong. Off one, you are not stranded: the documentation covers it, and I train your staff to handle the day to day themselves, the same way I did here. Questions are always welcome either way.

Retainer supportTrained staffDocs that stand alone

You bring the process, I map it

You do not need to arrive with a plan, a tool list, or any technical detail. Describe the manual work eating your time and I take it from there. Anything I need from you, accounts, access, a reference file, a person to check output, is listed in discovery and planning before a quote, so there are no surprises later.

Just describe the problemYou approve the price before I build

Available for workflow engineering and data operations roles.

Open to automation projects, contract work, and full-time positions.