Exploring O*NET's graph structure with TuckER to find related occupations
If you've ever looked into changing careers, you've probably run into a version of this question: which occupations match my current skills? One of the richest resources for answering it is O*NET, the Occupational Information Network. It describes 1,016 occupations through detailed profiles that include skills, knowledge areas, abilities, tasks, and work activities.
The standard approach to finding related occupations is straightforward: look at shared elements, and if two occupations have a lot of the same skills and knowledge areas, they're probably related. This works, as long as the data is available. However, could there be other ways to find occupations that are less obviously related?
O*NET publishes its data in RDF — a graph format — and that allows us to approach the problem from a different angle. Instead of comparing occupation profiles element by element, you can treat the entire dataset as a knowledge graph and look for patterns in its structure. This lets you identify related occupations even when they don't share many of the same skills, knowledge areas, or abilities.
I extracted a subgraph from O*NET 30.3 — a network where occupations, skills, knowledge areas, abilities, tasks, and detailed work activities (DWAs) are all nodes, connected by typed edges. Here's what a single occupation looks like in that graph:
Each occupation connects to its skills, knowledge areas, abilities, tasks, DWAs, job zone, and other related occupations. O*NET has a lot more information than what I've included here (work styles, interests, work context, education, wages, and more), but I kept the scope to the elements most directly tied to what a person does and knows in a role. I also filtered out low-importance elements using O*NET's own importance and level ratings, so the graph focuses on the strongest signals.
Out of the 1,016 occupations in O*NET, 923 ended up in the graph. The 93 that didn't make it are mostly "All Other" residual categories and military specializations — occupations without detailed element profiles.
For finding related career paths, one can simply compare element profiles directly. You may have noticed in the visualization above that O*NET already catalogs related occupations, which they determine largely by doing such a comparison (although, it should be noted that they also rely on judgment from human experts). That works well, but it can only find occupations that look alike on paper or that consulting experts can identify. Two roles might be structurally related — playing similar roles in the broader occupational landscape — without sharing many of the same elements.
Since O*NET's data is already structured as a knowledge graph with typed relationships, there's another angle worth exploring: models that learn from the pattern of connections, not just which elements are shared. To that end, I used a model called TuckER, which learns a low-dimensional representation of every entity and relationship in the graph and then uses those representations to predict missing links.
TuckER works by decomposing the full graph into a small core tensor and two embedding matrices — one for entities, one for relations — with far fewer dimensions than the original graph. This compression forces the model to learn the underlying patterns that connect occupations, skills, knowledge, and other elements, rather than memorizing individual links. It can then use those learned representations to score how likely any unobserved link is.
I trained the model on the O*NET knowledge graph and ranked all unobserved occupation pairs by predicted score, filtering out every pair that O*NET already lists as related. Here are a few examples of the top predictions that result:
The strongest cluster is in law enforcement and public safety. Police supervisors get paired with customs officers, fraud examiners, security managers, and fire investigators — all roles that share investigative and supervisory patterns, even though O*NET doesn't list them as related. Management analysts linked to HR specialists and forest fire inspectors linked to environmental compliance inspectors also seem like plausible fits.
The results are mixed, though. Judges, magistrate judges, and magistrates are also in the same law enforcement cluster, linked to first-line supervisors of police and detectives. Clearly, structural similarity in the graph doesn't always translate to a practical career move.
Here's a table of the top 25 predictions found by TuckER, along with their Jaccard similarity:
Pairs are ranked by TuckER's raw link-prediction score. Jaccard similarity (0–1) measures the overlap in skills, knowledge, abilities, tasks, and DWAs between the two occupations.
| # | Source | Target | Score | Jaccard |
|---|---|---|---|---|
| 1 | First-Line Supervisors of Police and Detectives | Customs and Border Protection Officers | 11.20 | 0.391 |
| 2 | First-Line Supervisors of Police and Detectives | Fraud Examiners, Investigators and Analysts | 10.80 | 0.333 |
| 3 | First-Line Supervisors of Police and Detectives | Judges, Magistrate Judges, and Magistrates | 10.72 | 0.347 |
| 4 | First-Line Supervisors of Police and Detectives | Airfield Operations Specialists | 10.65 | 0.314 |
| 5 | First-Line Supervisors of Police and Detectives | Security Managers | 10.59 | 0.364 |
| 6 | First-Line Supervisors of Police and Detectives | First-Line Supervisors of Personal Service Workers | 10.37 | 0.301 |
| 7 | Compensation, Benefits, and Job Analysis Specialists | Treasurers and Controllers | 10.30 | 0.321 |
| 8 | First-Line Supervisors of Police and Detectives | Management Analysts | 10.23 | 0.397 |
| 9 | Management Analysts | Human Resources Specialists | 10.22 | 0.338 |
| 10 | Concierges | Cashiers | 10.22 | 0.050 |
| 11 | Railroad Conductors and Yardmasters | Commercial Pilots | 10.19 | 0.244 |
| 12 | Police and Sheriff's Patrol Officers | Arbitrators, Mediators, and Conciliators | 10.18 | 0.215 |
| 13 | Management Analysts | Clinical Data Managers | 10.17 | 0.384 |
| 14 | Legal Secretaries and Administrative Assistants | Title Examiners, Abstractors, and Searchers | 10.17 | 0.286 |
| 15 | First-Line Supervisors of Security Workers | Airfield Operations Specialists | 10.16 | 0.333 |
| 16 | Forest Fire Inspectors and Prevention Specialists | Environmental Compliance Inspectors | 10.15 | 0.277 |
| 17 | Dispatchers, Except Police, Fire, and Ambulance | Air Traffic Controllers | 10.07 | 0.263 |
| 18 | Bailiffs | Compliance Officers | 10.03 | 0.220 |
| 19 | First-Line Supervisors of Police and Detectives | Government Property Inspectors and Investigators | 10.03 | 0.368 |
| 20 | Dispatchers, Except Police, Fire, and Ambulance | Freight Forwarders | 10.02 | 0.273 |
| 21 | Dispatchers, Except Police, Fire, and Ambulance | Airfield Operations Specialists | 10.02 | 0.306 |
| 22 | Legal Secretaries and Administrative Assistants | Statistical Assistants | 10.02 | 0.263 |
| 23 | First-Line Supervisors of Police and Detectives | Eligibility Interviewers, Government Programs | 10.01 | 0.297 |
| 24 | First-Line Supervisors of Police and Detectives | Fire Inspectors and Investigators | 10.00 | 0.302 |
| 25 | Management Analysts | Securities, Commodities, and Financial Services Sales Agents | 9.99 | 0.346 |
Element-based similarity methods find occupations that look alike on paper — shared skills, shared knowledge, shared abilities. TuckER tries a different angle: learning from the pattern of connections rather than counting shared elements. When it works, it surfaces connections like fraud examiners and police supervisors that share investigative reasoning patterns a skills checklist might not capture. When it doesn't, you can end up with things like railroad conductors paired with commercial pilots — superficially similar in the graph but not in practice.
These predictions are hypotheses, not recommendations, and the mixed results suggest this approach works better as a starting point. The natural next step would be validation against real-world data. The Bureau of Labor Statistics publishes occupational transition data that could tell us whether these predicted paths correspond to career moves people actually make.
TuckER is also just one approach. Other knowledge graph embedding models — RotatE, TransR, ConvE — model relational patterns differently and might surface different connections. Graph neural network methods like R-GCN could go further by incorporating node features and multi-hop reasoning.
TuckER uses Tucker decomposition to model relationships in a knowledge graph. For a given subject entity and relation , it scores all possible object entities:
where is a learnable core tensor, is the subject entity embedding, is the relation embedding, and is the full entity embedding matrix.
| Parameter | Value |
|---|---|
| Entity embedding dim | 200 |
| Relation embedding dim | 30 |
| Batch size | 128 |
| Learning rate (initial) | 0.005 |
| LR decay (per epoch) | 0.995 |
| Label smoothing | 0.1 |
| Dropout (input / hidden₁ / hidden₂) | 0.2 / 0.2 / 0.3 |
| Loss function | Softmax cross-entropy |
| Optimizer | Adam |
| Max epochs | 500 |
| Actual epochs (early stopping) | 237 |
| Patience | 50 |
Filtered link prediction metrics with 95% bootstrap confidence intervals (n = 1,000 resamples).
The original TuckER paper uses binary cross-entropy (BCE) loss. That didn't work here. With O*NET's extreme class imbalance — roughly 3 positive entities out of 21,900 per training example (0.014%) — BCE lets the model get away with predicting all-negative scores and still achieving near-zero loss. The model basically learns to say "nothing is related to anything," which isn't very useful.
Softmax cross-entropy fixes this by normalizing scores across all entities before computing the loss, which forces the model to actually rank positive entities above negatives. There's no trivial all-negative solution. Label smoothing () keeps the model from getting overconfident about its predictions.