HiBob lists twelve named agents on its AI page. Performance, Learning, Development, Skills, Talent, Finance, Surveys, Social, Hiring, Attendance, Payroll, Compensation.
If you turn all twelve on in the same week, you will have no idea which ones worked.
You'll have a spike in usage, a handful of anecdotes, one manager who loves it and one who doesn't, and no way to answer the only question that matters six months later: is this making anything measurably better, and is any of it making something quietly worse?
Feature-by-feature isn't caution. It's the only way to get an answer.
What you're actually enabling
Worth being precise about the naming, because three layers of it are in circulation and they don't line up cleanly.
Bob AI is HiBob's umbrella brand. Bob Companion is the conversational interface – the natural-language layer that routes a question to whichever capability can answer it. The twelve agents are HiBob's taxonomy for what sits behind that interface.
Here is the part that matters operationally: those twelve agent names are a marketing abstraction over a set of discrete features that shipped individually. Inside the product, you will find the feature names, not the agent names. Based on HiBob's own module pages, the names you'll be looking for include Writing Assistance, 360° Performance Review Summary, Goals and Key Results Generator and Theme & Sentiment Analysis on Talent; CV Summary, Scorecard Write-up, Evaluation Overview, Job Description Generator and Email Composer on Hiring; Budget Simulation, Pay Equity Analyzer, Comp Recommendations and Cycle Insights on Compensation; and AI Planning Assistant and AI-Built Scenario Modeling on Workforce Planning. Confirm the current list against your own tenant – HiBob has been shipping quickly and the marketing pages and the product don't always agree.
Plan the rollout against the feature list. It's what your admins will actually see and toggle.
One more status point, stated plainly because HiBob hasn't: the GA-versus-beta status of Bob Companion is unresolved. HiBob's pricing page lists "Bob AI Companion" as a Core plan feature, which reads as generally available. A third-party review in April 2026 described it as in beta testing. HiBob's subscription terms contain a Beta Services clause providing such features as-is and discontinuable without notice. Get the status of your tenant in writing from your account team before you build a rollout plan on top of it.
The constraint nobody mentions: your modules decide your agents
Bob AI is bundled into Core rather than sold as a separate per-seat SKU. That sounds like everything is available. It isn't.
The agents live inside modules. If you own Core only, your realistic starting inventory is Companion Q&A over core HR data (people, time off, documents, tasks), natural-language reporting, and recognition drafting. The Performance and Surveys agents need Talent. The Hiring agents need Hiring. Comp Recommendations and the Pay Equity Analyzer need Compensation. The Payroll Agent needs Payroll Hub.
So the first move in "which agents first" isn't a judgment call. Write down which modules you own. That eliminates most of the list before you've had a single conversation about sequencing.
For companies in the 15–50 range, be honest about the second half of this: several of these agents have no work to do. Calibration cycles, scenario planning, and skills frameworks are real problems at 300 people. At 35, they're a solution looking for a headcount.
The permission fact that makes staged rollout safe – and the one exception
HiBob is unusually direct about how Bob AI handles permissions:
"Bob AI respects the same permission layers as the rest of Bob. It cannot access or surface any data you are not authorized to view."
More usefully, HiBob's team has described the mechanism rather than just the outcome. In a published demo write-up from April 2026, HiBob described permissions as enforced at the query level, not by filtering after retrieval – the permission profile determines what data is included in the query in the first place, rather than the model retrieving broadly and then redacting. The reporter watched a request for the CEO's salary return nothing from a non-admin account and return the number from an admin account.
That is a vendor demonstration, not an audit. Reproduce it yourself on your own data before you rely on it.
That's the right design, and it's the reason feature-by-feature rollout is workable: turning on an agent doesn't widen anyone's access, it just makes existing access easier to use.
The exception is HiBob's MCP server, and it's a real one. MCP lets external AI clients – Claude, Cursor, VS Code, Copilot Studio, and the Slackbot integration launched in June 2026 – query Bob. At launch in April 2026 it authenticated with a service user, meaning everything the AI could reach was governed by that one service user's permission group, not by the permissions of the person asking. HiBob moved to OAuth in mid-2026, and its changelog is explicit: "AI clients connect on behalf of the signed-in Bob user, and access is limited to that user's permissions in Bob."
Two things follow. First, HiBob's own documentation was internally inconsistent on this point as of July 2026 – the integration guide still described service-user auth while the changelog announced OAuth. Verify what your tenant is actually doing. Second, and more important: audit your service users before you touch MCP or Slackbot. Service users default to no access, which is the right default, but the ones that exist were scoped for integrations by people who wanted them to work, not by people thinking about a conversational layer. That scope is the AI's blast radius.
This is also the reason we put a permission audit before an AI rollout rather than alongside it.
The order
Five waves. Each one has a reason it sits where it sits, a signal that tells you it's working, and a signal that tells you to stop.
Wave 1 – HR admin reporting
Turn on: natural-language reporting and the Companion interface, for the HR/People team only.
Why first: smallest blast radius. The audience is three to six people who already have broad access, so nothing new becomes visible. The output is checkable – if it returns a wrong number, someone notices immediately, because they know what the right number is. And it produces the fastest measurable win, because "build me a report" is the single most common interruption an HR ops person absorbs.
Success signal: ad-hoc report requests routed to HR ops drop month over month. Count them for four weeks before you enable anything – you cannot show a reduction against a baseline you never took.
Failure signal: queries that return nothing, return the wrong figure, or get quietly re-run in the classic report builder. Track the re-run rate. It is the honest measure of whether anyone trusts it.
Note the counter-evidence. Public reviews of HiBob's AI reporting are mixed. One reviewer wrote in late 2025 that AI "does not work at all for reporting – never gives me what I need." Another in March 2026 said the AI features were "not quite where I would like them to be." Pilot this. Don't announce it.
Wave 2 – Drafting agents
Turn on: Job Description Generator, Writing Assistance, Email Composer, and recognition drafting.
Why second: these are pure draft-not-decide. A human writes over the top of everything they produce, nothing is committed without a person clicking, and no sensitive data is exposed that wasn't already. High perceived value, near-zero governance surface.
Success signal: share of new requisitions that start from a generated JD; share of managers who use writing assistance more than once. Adoption is the metric here, because the value is time, not accuracy.
Failure signal: homogenization. HiBob's own AI terms warn that output "may not be unique and could be identical or substantially similar to output generated for other customers." That's a vendor telling you your job postings may start sounding like everyone else's. Read ten generated JDs side by side at week four. If they're interchangeable, the feature is saving time and costing differentiation, and that's a trade you should make deliberately.
Wave 3 – Employee-facing policy and time-off assistant
Turn on: Companion for all employees, answering policy, PTO, and benefits questions.
Why third, and why the gap: this is the agent employees actually see. It answers questions with financial and legal consequences, in a voice that sounds like the company's. And its grounding mechanism is the least-documented part of the product. HiBob says answers are "grounded in company data and policies." It does not publish which document stores are indexed, whether document-level permissions are honored during retrieval, or whether answers carry citations back to source.
Do the document work first. Curate a small, current, authoritative policy set. Archive superseded versions. Do not point it at the drive and hope. The best public benchmark on retrieval over internal company documents – an academic evaluation published in May 2026 across roughly half a million synthetic internal documents – put baseline correctness in the 51–69% range depending on retrieval method. Even good retrieval over messy internal content gets a meaningful share of questions wrong. Your corpus quality is the variable you control.
Success signal: percentage of employees who use it at least once a month, and inbound HR tickets per 100 employees.
Failure signal: the spot-check. Sample twenty answers a week against the source policy for the first eight weeks. Log every discrepancy. Also watch for escalations that follow a Companion answer – those are the expensive ones, because the employee acted on it.
One trust consideration. Survey work published in late 2025 found only 27% of US workers fully trust their employers to use AI responsibly, and more than half prefer a human to evaluate their performance. Announce clearly what this assistant does and, more importantly, what it does not decide. The announcement is part of the rollout, not an afterthought.
Wave 4 – Surveys and performance summarization
Turn on: Theme & Sentiment Analysis, 360° Performance Review Summary, Goals and Key Results Generator.
Why fourth: now you're touching subjective employee content, and the output feeds decisions about people. Still human-reviewed, but the stakes have changed.
Success signal: time from survey close to a readout leadership actually reads. Median manager hours per review cycle.
Failure signal: two things. First, analyst disagreement – have a person theme a sample of survey responses independently and compare. Second, review homogenization: measure text similarity across reviews written by the same manager, and track the share of AI drafts submitted with no edits. A manager who submits unedited drafts for their whole team has outsourced the part of the job that isn't delegable.
Wave 5 – Compensation and hiring
Turn on: Comp Recommendations, Budget Simulation, Pay Equity Analyzer, Cycle Insights; CV Summary, Scorecard Write-up, Evaluation Overview.
Why last: highest regulatory exposure and the most sensitive data category in the system. These get a legal gate, not just a change-management plan.
The gate:
- If you hire in New York City, candidate screening tools may fall under Local Law 144, which requires an independent bias audit within one year of use, a public summary of the results, and candidate notice ten business days before use. NYC is currently the only US jurisdiction that actually mandates a bias audit. HiBob publishes no bias-audit documentation for these features. That obligation is yours.
- If you hire in California, the FEHA automated-decision-system regulations effective October 1, 2025 reach tools that screen, score, rank, or recommend candidates even where a human decides. California does not require testing – but it makes the presence, quality, recency, and results of testing evidence in a discrimination claim. And the CCPA's automated decisionmaking rules bring compensation squarely into scope from January 1, 2027.
- If you hire in Illinois, amendments to the Human Rights Act effective January 1, 2026 add notice obligations for AI used in recruiting, hiring, and promotion.
We wrote the policy side of this up separately in HR AI policy for California employers.
Success signal: for compensation, cycle completion time and the share of manager proposals that start from a recommendation.
Failure signal: override rate, and – the one that matters – whether new pay gaps opened during the cycle. AI trained on historical pay data can replicate historical pay disparities with more confidence and better formatting. Run the equity analysis after the cycle, not just before it.
Off to one side: MCP and Slackbot
Not a wave. A separate track with its own prerequisites: confirm your tenant is on OAuth rather than service-user auth, audit every service user's permission group, and decide explicitly whether people-data questions should be answerable inside Slack. HiBob's framing is fair – Slack isn't exposing new data, it's surfacing approved data to people already authorized to see it – but "already authorized" is doing a lot of work in that sentence, and you're the one who set the authorizations.
How to tell any of it is working
HiBob publishes essentially no efficacy metrics for Companion. No deflection rate, no time-saved figure, no accuracy benchmark, no adoption data. That absence is worth naming, and it means you have to instrument this yourself.
| Wave | Leading metric | Lagging metric | Failure metric |
|---|---|---|---|
| 1 – HR reporting | Reports built via natural language vs. builder | Ad-hoc report requests to HR ops | Queries re-run in the classic builder |
| 2 – Drafting | Share of reqs using a generated JD | Recruiter and manager hours per req | Similarity across generated output |
| 3 – Employee Q&A | Monthly active employee users | HR tickets per 100 employees | Spot-check error rate; post-answer escalations |
| 4 – Surveys / performance | Time from survey close to readout | Manager hours per review cycle | Analyst disagreement; unedited-draft rate |
| 5 – Comp / hiring | Share of proposals opened from a recommendation | Cycle completion time | Override rate; new pay gaps; adverse-impact drift |
Take the baseline before wave one. Every number in the middle column is meaningless without one.
Five things to get in writing before you start
- Is Companion GA or beta in our tenant? The Beta Services clause in HiBob's subscription terms provides such features as-is and allows discontinuation without notice.
- An ISO 42001 certificate, if one exists. Read HiBob's AI trust page carefully: it says the company "follows industry best practices, including ISO 27001, ISO27018, ISO42001." That is a claim about alignment, not a claim to hold a certificate, and it's a materially weaker statement than it looks at a glance. It also doesn't appear on HiBob's main security page. Rippling, by contrast, lists ISO 42001 among its certifications. Ask HiBob directly whether a certificate exists and, if so, for a copy and the certifying body.
- Where does model inference actually happen? HiBob's sub-processor list names OpenAI, Microsoft (Azure OpenAI), and AWS EMEA in Luxembourg as AI sub-processors, and gives the processing location for all three as the EU, with standard contractual clauses and data privacy framework adequacy as the transfer mechanisms for the two US-incorporated vendors. HiBob also says customer data sent to AI providers is not stored or logged, and its AI terms state customer data is not used to train models. What it does not publish is the specific region in which inference executes, or how that is guaranteed. For a client with EU staff, that is the question – and confirm the platform hosting regions separately rather than relying on the commonly repeated "Ireland with a German backup."
- Is AI usage logged, and can we export it? We found no mention of AI activity logging in HiBob's published AI governance material. "Who asked Companion what" is the audit record you'll want the first time something goes wrong.
- Does AI stay bundled into Core? There is no published per-seat AI licence, no message metering, and no consumption cap. "Included" today is not a contractual guarantee of "included" at renewal. Put it in the order form.
The order, compressed
- HR admin reporting – smallest audience, checkable output. Take the ticket baseline first.
- Drafting agents – JDs, review writing, email, recognition. Watch for homogenization at week four.
- Employee policy and PTO Q&A – only after document curation, with an eight-week spot-check.
- Surveys and performance summarization – add the analyst-disagreement and unedited-draft checks.
- Compensation and hiring – legal gate first: LL 144, FEHA ADS, Illinois notice. Re-run pay equity after the cycle.
- MCP and Slackbot – separate track: confirm OAuth, audit every service user, decide whether people data belongs in Slack at all.
And before wave one: confirm GA-vs-beta in writing, and get the five contract answers below.
People Street's take
The instinct with a platform like Bob is to switch everything on and let people find what's useful. It feels generous. It's actually the fastest way to end up with a company that has AI everywhere and evidence nowhere.
Feature-by-feature costs you a few weeks. What you get for those weeks is a rollout you can explain – to a nervous employee, to a board that asked what the AI investment returned, and to a regulator asking who reviewed the output before it affected somebody's pay.
Start with the agent whose failure is visible and whose audience is small. Save the ones that touch compensation and candidates until you have a legal answer, not a vendor assurance. And take the baseline first, because the whole exercise is worthless if you can't say what changed.
If you're on Bob and planning an AI rollout – or you turned it all on already and want to know what you're exposed to – book a 20-minute call.
Related: The permission audit your AI needs · Rippling AI vs Bob Companion · Rippling vs HiBob