Six people sit in a meeting room.
A VP.
A senior engineer with twenty years in the work.
A young AI engineer who has only been at the company two years.
A sales lead.
A finance lead.
And a product manager.
They have to decide something very concrete:
an important project — should it continue?
The money has already been spent. Time has already gone. There is competition outside the room, and risk inside it.
So a question that looks simple appears:
Of these six people, whose words should weigh more?
The traditional company has a very cheap answer.
Look at the title.
The VP’s words weigh most. Then seniority. Then the people who own the business. The person who arrived two years ago usually sits at the back.
Why?
Because it is the least trouble. The org chart is already drawn. Whoever sits higher already holds more power.
The last essay asked why a company’s authority should belong, for the long term, to a title.
Open that door, and the next question walks in immediately.
Fine.
If a title should not decide forever who gets the last word —
then who should we actually listen to?
The most natural answer may be: listen to the experts. Listen to the people who know most.
It sounds far more reasonable than listening to the highest title.
Ask one more question, though:
Who defines the expert?
The problem is immediately messy again. Twenty years in the job? How many projects? An elite degree? Last year’s performance rating? How many past wins? Or should AI compute a score for each person from historical data?
The senior engineer will feel: I have seen the most failures.
The young engineer will feel: I am closest to the newest technology.
Sales will say: only I know what customers actually think.
Finance will say: you are all telling stories. In the end the numbers speak.
The product manager will believe: I know best why users pay.
The VP has a reason too: I see the whole company, not one local problem.
Each of them can prove they should be the person in the room most worth hearing.
So “listen to the experts” does not really solve the problem. It only swaps the title for another identity — seniority, résumé, degree, past results. Still identity.
Just another costume.
Then it is easy to take one more step. People are biased. Let the system calculate.
Jason: 87.4.
Alice: 76.2.
Bob: 61.8.
Whoever scores higher, their words weigh more. Whoever scores lower listens first.
It looks scientific. It even looks far fairer than a title, because it no longer asks whether you are a VP. It asks how you actually performed.
That is also the problem.
If we are honest enough, we will notice: swap Title for Score, and we may have changed nothing.
Before: VP > Director > Manager.
After: 92 > 84 > 71.
The hierarchy is still there. It has only moved off the badge and into the database.
It may even be more dangerous. The old title system at least admitted: this is power. An algorithmic score puts on a prettier coat.
This is not power.
This is science.
Once the system writes someone as 61.8, they may already walk into the room a little shorter — even if today’s question is exactly the one they know best.
So I more and more think the question we should actually ask is not: is this person a strong judge?
It is: on today’s concrete problem, why is this judgment worth taking seriously?
Those two questions sit very far apart. The first scores the person. The second evaluates this one judgment.
Someone may be excellent at judging why a product will fail, and still not know whether consumers will like a new flavor. They may read supplier mold risk with startling accuracy, and on the macroeconomy they may only be reading the news. They may be very good at engineering trade-offs, and on regulatory policy they may still be guessing.
A person should not have to carry a universal Judgment Score, like a credit score, into every problem they walk into.
That is also a principle I think future companies will need:
Score this judgment. Do not give a person a total score.
Score predictions, not people.
What does it mean, then, to score a judgment?
Start with a very simple change. Do not only ask the company: will this project succeed?
That question is too easy to fake knowing. Yes. No. It depends. Anyone can say those.
Ask it differently: what is the probability this project successfully ships before December 1?
Alice: 70%.
Bob: 40%.
Jason: 55%.
Suddenly the three judgments can be recorded. What is worth watching is no longer who speaks with more force, or who holds the higher title. It is this:
When Alice says 70%, do things of this kind later happen about seven times in ten?
When Bob keeps saying 40%, is he really seeing the risk earlier than everyone else — or is he habitually pessimistic?
That is a concept that matters a great deal in forecasting research: Calibration.
It is not asking whether you are smart. It is asking whether the confidence you claim matches what later happened.
Guessing right three times in a row may only be luck. But if someone has made many similar judgments, and the things they call 70% really do happen about seven times in ten over the long run, and the things they call 40% happen about four times in ten — it starts to look less like luck. It is closer to a capacity for judgment that has been checked.
Notice, though. This still does not mean Alice is an 87-point person. It only means: on this class of problem, Alice’s forecasts are worth a more serious look.
The problems a company actually faces are rarely this clean. Many decisions take a year before anyone knows the result. Some never get a clear right or wrong. Some problems human beings have never met.
So if someone tells you we can already calculate every employee’s judgment ability with precision, I become more wary. An honest system often will not hand you a pretty score. It will keep asking a few plain questions.
How many similar problems has this person actually handled? Not: how many years have they been at the company. But: have they handled a similar supplier incident? Similar regulatory risk? A similar product launch?
And where does the judgment come from? Directly measured data? System records? Cases that keep returning? Firsthand experience? Someone else’s retelling? Or an intuition they can barely explain, formed only after doing the work many times?
I think one point here is especially important. Intuition should not automatically be treated as low-quality evidence. People who have actually done this for decades form a kind of judgment that is very hard to put on a slide. They cannot write the full formula. They just feel: this does not smell right.
That should enter the system too. The system only needs to know what kind of evidence this judgment sits on — not pretend, for the sake of looking scientific, that every judgment is the same kind of thing.
There is a more practical problem as well.
Is this sentence even their own judgment?
Suppose the VP opens: I personally think this project has at least an 80% chance of success.
Then the other five start talking.
82%.
75%.
78%.
80%.
85%.
Are those really six judgments?
Not necessarily. They may only be one judgment, plus five echoes.
So I think the decisions that will actually matter in the future should perhaps begin with something many companies are not used to doing.
Do not speak first. Each person independently writes their judgment. Lock it. Then discuss. Exchange evidence. Hear the pushback. Then allow a revision.
The company can leave four things behind: the initial judgment, the new evidence, the revised judgment, and the final result.
If none of that is left, what the company remembers years later is usually only what the most powerful person in the room said at the time.
And something else in the meeting is especially easy to erase.
The minority view.
Suppose six people judge the probability the project succeeds. Five of them say:
75%.
78%.
80%.
72%.
77%.
Alice says:
30%.
What does a traditional meeting likely do? Five people are roughly aligned. Alice is more conservative. Continue.
If we let AI compute a simple weighted average, the situation may not be much better. What appears is 68.7%. It looks extremely precise. People may even feel the system has already figured it out.
The valuable thing may have been washed out by that average.
Because AI can keep looking. Alice has handled five very similar failures. On this class of problem, her past forecasts were not random shouting. She gave 30% before hearing anyone else’s answers. And she pointed to a regulatory risk the other five never mentioned.
At that point, the valuable AI output should not only be Final Probability: 68.7%. It should be: there is a disagreement here worth reopening.
That is a principle I think future judgment systems will need. The point of aggregation is not only to average opinions into a number. What matters is to surface the disagreement that is worth seeing.
Where do people actually agree? What is the real disagreement? Is there a minority view sitting on experience or evidence the others do not have? Whose past work is actually relevant to today’s problem? And what do we, in fact, not know?
Those questions are often worth more than a mysterious total score. Because what a company should actually fear is never that someone in the room disagrees.
It is that the “difference” most worth hearing gets averaged away, politeness-smoothed, and title-covered — lightly.
Here we have to draw a very clear line.
Even if Alice’s forecast deserves serious weight, that does not mean Alice automatically holds the final decision.
Prediction is not Decision.
Prediction asks: what is most likely to happen? Decision asks: knowing those possibilities, what do we actually do?
A great deal sits in between. Strategy. Risk appetite. Law. Values. Cash. Opportunity cost. And the most important thing: if this is wrong, who catches the result?
A team can correctly judge that the chance of shipping on time is only 32%. That does not automatically mean: cancel. The company may want to buy that 32%. It may be the only window still open. Not doing it may cost more. The company may even know the success rate is low and still have to do it for strategy.
So a judgment system can influence authority. It cannot automatically become authority.
Knowing what might happen, and deciding which future we are willing to carry, remain two different things.
If the essay ended here, the system might already sound beautiful.
Of course it is not that simple. It gets sick too.
The first cut lands on newcomers.
Suppose historical judgment starts to affect whether someone is worth hearing now. What happens to the person who just joined? They have no record. No history. The system does not know whether they are good.
There has to be a very important principle here:
Unknown is not Bad.
No history does not mean low ability. Real organizations make this mistake easily. No record becomes a low score, because a blank is what least gives leaders a sense of safety. Young people, or anyone new, then sit at the back by nature — not because they were wrong, but because the system has not seen them yet.
The future may let newcomers first build a judgment record on low-risk problems with fast feedback. It may let them make shadow forecasts: leave their judgment, but do not let it affect the formal decision yet. Those things can still be designed.
They are not the answer yet. This essay only needs to block one move:
Never write “not yet verified” as “not worth hearing.”
The second cut is gaming.
Once a judgment record starts to affect someone’s influence, people will start managing that record. That is human nature.
They will pick the easy-to-forecast questions. They will dodge the decisions with real uncertainty. They will submit a little later. They will see what others think first. They will hide how unsure they are. After the fact they will reinterpret what they meant. They will make the score look good, instead of making the company better.
That is not necessarily because employees are bad. It is because the system started handing out rewards. Once a metric moves from helping you understand to deciding someone’s future, people start living around the metric.
So this distinction has to stay in place:
Measurement.
And:
Governance.
You can record someone’s forecasts. That does not mean you already have the right to decide their career from that record.
You can design some protections. Lock the initial judgment first. Look at different domains separately. Let old records lose weight. Allow challenge. Distinguish people who will touch hard problems from people who only farm easy ones.
None of that kills gaming. It only makes gaming harder. Any system that tells you it has completely solved employee score-farming is the one most worth watching.
The third cut may be the most dangerous.
Suppose Alice really has been accurate for a long time. So the organization listens to her more. The more it listens, the easier it is for her to enter more important projects. More important projects bring more information, resources, and experience. Then her record grows stronger still.
What happens in the end?
We may have remade a permanent VP. Only this time the title is gone, and the hierarchy lives in the database.
That is the result I think the whole judgment system has to prevent. We cannot kill the permanent VP and then remake a permanent VP in the database.
People who already have recognition find it easier to keep receiving it. That is not a new phenomenon. Companies work this way. Academia works this way. Social platforms work this way too.
So if something called judgment weight really exists, it at least has to be about this problem, about this domain, able to expire, able to be challenged, and able to admit: this time, I do not know.
It must never become a new permanent identity.
Then there is a fourth cut. In the AI era, it is an especially messy one.
New problems.
Historical calibration fails most easily exactly when the company most needs judgment: things nobody has seen. The AI era will produce more of these. They did not exist three years ago. Past records are not comparable. Past winners may no longer be accurate.
At that point the system has to be willing to say: We don’t know.
It cannot manufacture authority because history is missing. It cannot assume that because Alice was especially accurate on old problems, she should still weigh most on a brand-new one.
Sometimes the most honest weight is: this time, nobody has privilege.
So what should AI actually do here?
I now think the answer is getting clearer. First, what it should not do.
AI should not become a Management God. It should not tell the company that Alice should receive 37% of the power. Power is not calculated that way.
What AI is actually suited for is quite concrete. Find past decisions that were truly similar. Look at different people’s histories of probability judgment. Surface evidence that conflicts. Notice which opinions are only copies of each other. Mark the minority view. Keep the first judgment and the later revision. Point out what evidence is still missing. Clear the context of a complicated decision.
In other words, AI’s most important job may not be helping the company decide directly. It may first help the company decide where to look carefully now. Who is worth one more listen. Which risk should not be slid past. Which piece of evidence has not yet entered the room.
That is a distinction I like. AI can help allocate attention. It should not automatically allocate sovereignty.
AI can help authority find better judgment. Final formal authority should still land on people and rules that can be explained, challenged, and held to account.
Otherwise we have only outsourced one kind of judgment for another: hand the judgment to the model, and leave the responsibility to the air.
After the first round of research, I think five principles are still standing, for now.
They are not laws. They are not truths Future Lab has already proven. So far, we have not punched through them ourselves.
First:
Score this judgment. Do not give a person a total score.
Score predictions, not people.
Second:
Weight this opinion. Do not permanently weight an identity.
Weight judgments, not identities.
Third:
Record the contribution. Do not rank a person’s worth.
Record contribution, don’t rank human worth.
Fourth:
Surface the disagreement that is worth seeing. Do not average it away.
Surface disagreement, don’t average it away.
Fifth:
Let AI help authority find better judgment. Do not let AI itself become the source of authority.
Let AI inform authority, never automatically become authority.
The easiest one to get wrong is the first. Scoring the judgment slides too easily into scoring the person. The system will want a leaderboard. Leaders will want a list. Employees will want a number they can farm.
In the end, prediction becomes identity again.
So this essay still has not solved the future company. It has only answered the question the last essay left behind.
If a title should not permanently decide who gets the last word, we probably should not immediately find another way to rank people.
What we should actually do first is learn to judge: on today’s concrete problem, which judgment, and why, is worth taking seriously?
But the moment we start recording those judgments, the next door appears.
Suppose we really want to know what information this person had at the time. What they believed. What probability they gave. Why they judged that way. Whether they later changed their mind because of new evidence. And what finally happened.
Then the company has to solve another problem: how should a company remember a decision?
Once the result is out, memory starts quietly rewriting history. On the successful projects, everyone remembers they were supportive. On the failed ones, everyone remembers they “had reservations early.”
If a person’s judgment record will actually affect who is more worth hearing, the organization has to leave that moment’s judgment behind before the result appears.
Otherwise so-called historical calibration is still only after-the-fact writing.
And after-the-fact writing has always listened more to power.
Appendix: Key Concepts
- Score Predictions, Not People
- Evaluate a person’s judgment on a concrete problem. Do not build a permanently valid ability score that follows them across every scene.
- Judgment Weight ≠ Decision Authority
- That a judgment deserves serious weight does not mean the person who made it automatically holds the final signature.
- Prediction ≠ Decision
- Between what is likely to happen and what we decide to do sit strategy, risk, law, capital, values, and accountability.
- Blind First
- On important decisions, participants first leave their initial judgments separately, then discuss, exchange evidence, and update — so titles and the group contaminate independent judgment less.
- Unknown ≠ Bad
- No historical record means unknown. It does not mean low ability.
- Attention ≠ Sovereignty
- AI can help an organization see which opinions and evidence deserve more attention. That should not automatically confer formal decision rights.
- Decision Snapshot
- Before the final result appears, save what was known at the time: the information, the judgments, the probabilities, the evidence, and later updates. The next essay continues this.
Research Foundations
This essay borrows existing work in forecasting science, structured expert judgment, collective intelligence, and algorithmic management to test and constrain the argument. It does not rebrand those theories as original Future Lab concepts.
- Brier (1950); Murphy & Winkler — probability forecasts and calibration
Probability judgments can be tested over time. When someone says “70%,” whether events of that kind later happen about seven times in ten is more meaningful than simply counting how often they “got it right.” - Philip Tetlock — Expert Political Judgment; Good Judgment Project
Expert identity itself does not guarantee forecast quality. Probability calibration, continuous updating, and revising a judgment when the evidence turns are important parts of high-quality forecasting. - Cooke Classical Model; structured expert judgment research
In some settings, experts can be given different judgment weights based on how they performed on related calibration questions. Complex weighting is not always better than a simple average. - IDEA Protocol / Delphi-style structured judgment methods
Judge independently first, then discuss, update, and aggregate. That can reduce the information lost when a group starts influencing itself. - Prediction Markets — Google, HP, Intel and other company practice
Organizations can aggregate scattered information through internal prediction markets. Forecast accuracy still does not equal formal decision rights. - Lorenz et al.; research on the wisdom of crowds and social influence
Once individual judgments stop being independent, the group may converge quickly without becoming more accurate. - Prelec, Seung & McCoy — Surprisingly Popular
The majority view does not always hold the most valuable information. Some informed minority judgments are worth identifying on their own, rather than being averaged away. - Kahneman & Klein — expert intuition and feedback environments
In domains with long, stable, and clear feedback, experience can form valuable intuition. Knowledge that is hard to put into words should not automatically be treated as noise. - Goodhart’s Law / Campbell’s Law
When a metric starts deciding rewards and power, people adjust their behavior to optimize the metric itself, not the real thing the metric was meant to measure. - Merton — Matthew Effect
People who already have recognition find it easier to keep receiving opportunities and resources. Any weighting system based on historical performance can gradually grow a new permanent elite. - Kellogg, Valentine & Christin — Algorithmic Management
Algorithmic scoring and automated management do not naturally mean empowerment. They can also turn traditional management into a more hidden, more concentrated, harder-to-challenge form of control. - Fischhoff — Hindsight Bias
After people know the outcome, they systematically overestimate how much they “already knew.” Judgment records need to be saved before the result appears.