Developing and Validating a Leadership Assessment Based on The Inner Development Goals Framework
Project Summary
Our partner wanted a reliable leadership assessment tool that could measure leadership behaviours consistently while providing meaningful developmental feedback. Using the Inner Development Goals framework as the foundation, we designed a six-domain psychometric assessment and guided it through the full development process, from behavioural item writing and expert review to statistical validation. The project resulted in reliable long and short assessment forms ready for confirmatory validation and future organisational use.
The Challenge
We wanted to transform the Inner Development Goals framework into a scientifically robust leadership assessment that could be used confidently for leadership development. This meant translating broad leadership concepts into clearly defined behavioural constructs, designing assessment items that measured those behaviours consistently, and demonstrating that the assessment met accepted psychometric standards. At the same time, the tool needed to remain practical, engaging, and suitable for real-world organisational use.
Project Goals
- Develop a leadership assessment based on the Inner Development Goals framework.
- Define the assessment architecture across leadership domains, dimensions, and facets.
- Generate and refine an item pool through literature review and scale adaptation.
- Establish content validity through expert review and item representativeness analysis.
- Collect, screen, and prepare participant data for psychometric validation.
- Evaluate the underlying factor structure using exploratory factor analysis.
- Develop reliable long and short versions of the assessment.
- Design the scoring methodology and interpretation framework.
- Create narrative feedback content, visual reports, and leadership profiles.
- Develop the scoring algorithm and reporting workflow for platform integration.
Our Solution
We transformed the Inner Development Goals framework into a leadership assessment tool designed to measure six core leadership domains, alongside leadership confidence and positive self-presentation. Each domain was broken down into measurable dimensions, facets, and behavioural indicators, creating the foundation for item development. Drawing on leadership research and existing assessment literature, we developed an initial pool of 764 items focused on observable leadership behaviours rather than broad personality traits.
The items were rated by experts for representativeness using inter-rater agreement analysis. The assessment was refined, and the item pool was reduced to 117 items in 24 dimensions, which were then combined into scales for psychometric analysis, reliability and validity.
We then conducted a large-scale data collection phase with 352 professionals through an online survey. The sample represented a range of leadership experience levels, with the largest group (46%) reporting between three and five years of leadership experience. To ensure reliable analysis, we applied multiple data-quality procedures, including attention checks, Mahalanobis distance analysis, and response pattern screening, removing careless responses and statistical outliers and retaining 213 high-quality cases for psychometric evaluation. The majority of participants were based in the USA (97%). Before analysing the assessment structure, we confirmed the dataset was suitable for factor analysis through KMO and Bartlett's tests and screened for overlapping items through multicollinearity analysis.
We evaluated the assessment structure using exploratory factor analysis to identify the underlying leadership dimensions represented within the data. Using principal axis factoring with oblimin rotation, appropriate for correlated behavioural constructs and non-normal data, we identified an eight-factor structure with strong model fit indicators. We refined the assessment by removing weaker items and produced two versions: an 80-item long form and a 50-item short form. Both versions showed good reliability, with internal consistency reaching .89 for the short form and category reliability ranging from .68 to .82 for the long version.
We also evaluated evidence of convergent and discriminant validity using Average Variance Extracted (AVE), identifying areas with stronger measurement support alongside categories requiring further refinement during future confirmatory factor analysis (CFA).