Most HR teams already use some version of the 9-box grid. And it’s worth being honest about it: the tool is useful. It creates sharper talent conversations and forces organisations to differentiate rather than default to “everyone is doing fine.”
The problem isn’t the grid.
It’s the axis.
“Potential” is notoriously difficult to assess without a clear definition and evidence-based tools. And when the definition is vague, bias has plenty of room to enter.
This matters even more when deciding who gets a seat in a Women’s Leadership Development Programme (WLDP).
At Breath Beings, we recommend keeping the 9-box structure but changing what the axes measure:
Impact (x-axis): What has someone actually delivered? Outcomes, goal achievement and value created in their current role.
Growth Velocity (y-axis): How quickly is someone expanding their capability? How do they respond to stretch, absorb complexity, learn from feedback and take on more?
Impact tells you what she has done. Growth Velocity tells you how she is developing.
And that distinction matters because “potential” is one of the places where gender bias can quietly enter the talent process.
The Axis is Where the Bias Lives
Research has repeatedly shown that men and women can be assessed differently on performance and potential.
Danielle Li’s research, covering around 30,000 employees at a large retail organisation, found that women were rated higher on performance but lower on potential than men. The difference in potential ratings accounted for a substantial share of the organisation’s gender gap in promotions.
A meta-analysis by Player et al. (2019) similarly found that men are more likely to be selected for leadership based on perceived potential, while women are more likely to be expected to demonstrate proven performance.
The issue isn’t simply that “bias exists.”
It’s that the same behaviour can be interpreted differently.
Ambition in one person can look like leadership potential. In another, it can be labelled as being overly assertive. Confidence can become “executive presence” for one employee and “difficult behaviour” for another.
And the problem starts even before someone enters the 9-box conversation.
If managers are asked to identify who has “leadership potential,” they are often being asked to make a prediction based on a relatively small set of observations. Visibility, confidence, communication style, proximity to leadership and familiarity can all influence that judgement.
That makes potential a particularly slippery metric.
What Happens When Women have to Prove Potential?
There is another layer to this.
Women are often expected to demonstrate strong performance before they are considered ready for the next opportunity. The result can be a familiar cycle:
Perform → prove → perform some more → then be considered for potential.
Meanwhile, perceived potential can itself become a reason for giving someone the stretch opportunity that eventually creates the evidence of readiness.
This creates a catch-22.
You need experience to demonstrate readiness. But you often need to be perceived as ready to receive the experience in the first place.
Research by Exley and Kessler has also found gender differences in how people assess their own performance and potential. That matters when leadership programmes rely heavily on self-nomination.
Self-nomination can be useful. But it shouldn’t become the primary proxy for ambition or readiness.
Someone being comfortable saying “I am ready for the next level” is not necessarily the same as someone demonstrating the behaviours that indicate they can grow into it.
So What Should we Measure Instead?
This is where Growth Velocity becomes useful.
Growth Velocity cannot simply become “potential” with a new name. It needs to be observable and evidenced.
Instead of asking:
“Does she have leadership potential?”
ask:
How has she grown?
Has she successfully taken on increasingly complex work?
Has she moved beyond her existing expertise and learned quickly?
Does she actively use feedback to change her approach?
Has she closed capability gaps over time?
Has her scope, influence or decision-making increased?
Has she demonstrated the ability to navigate ambiguity, work across stakeholders or create impact beyond her immediate role?
These questions shift the conversation from personality to evidence.
They also make the assessment more useful.
Because leadership development isn’t about identifying people who already look like leaders. It is about identifying people who are demonstrating the capacity to grow into greater levels of responsibility and complexity.
The Selection Process Matters too
Even a better framework can fall short if the nomination process remains subjective.
Self-nomination alone can miss people who are less comfortable advocating for themselves.
Manager nomination alone can reproduce the very biases that the framework is trying to address.
A stronger approach is to triangulate self, peer and manager input against the same evidence-based questions.
For example:
- What is one example of this person successfully taking on greater complexity?
- What capability has she developed significantly in the last 12–18 months?
- How has she responded to feedback?
- Where has she demonstrated influence beyond her formal role?
- What is the next level of challenge she appears ready to take on?
The purpose isn’t to collect three opinions and average them out.
It is to look for patterns across perspectives.
And then comes calibration.
Because even evidence-based criteria can be interpreted differently by different managers.
One manager may describe someone as “highly strategic.”
Another may describe a similar behaviour as “taking too much ownership.”
One may see assertiveness as leadership.
Another may interpret it as being difficult.
Calibration creates a space to question those inconsistencies before they become selection decisions.
The Final Test
A simple question can tell you whether your WLDP selection framework is working:
Can the selection committee explain why every person is on the Growth Velocity axis using specific evidence?
If the answer is:
“She has a lot of potential.”
You probably haven’t defined the metric tightly enough.
If the answer is:
“She has successfully taken on larger projects, adapted quickly to new stakeholders, acted on feedback and consistently expanded the complexity of her work.”
Now you have something you can assess.
The goal isn’t to remove judgement from talent decisions. That’s probably impossible.
The goal is to make that judgement more consistent, observable and defensible.
Because selecting women for leadership development shouldn’t be about identifying the people who already look like leaders.
It should be about recognising the women who are already demonstrating growth, and making sure the process is designed to actually see it.