The RAIM Delphi Study
How much does each stakeholder group value each responsible-AI feature? Explore the results of our two-stage Delphi study.
To validate and prioritise the 45 candidate features of the RAIM framework, we ran a two-stage Delphi study with a panel of music AI researchers, artists & creators, AI ethicists, and legal experts. In Stage 1, panellists rated the importance of every feature (1–5 Likert scale) independently. Results were aggregated, shared back with the panel, and discussed; in Stage 2 the same panellists re-rated each feature, aiming to reduce disagreement and move towards consensus across groups. The charts below let you explore both what each group values (per-pillar radar charts) and how opinions shifted between the two stages.
Feature importance by stakeholder group
Each axis is a feature within the selected pillar; the distance from the centre is the mean importance score (1–5) given by that group in Stage 2. The dashed ring marks the neutral midpoint (3). Select a pillar below, and click a group in the legend to isolate it.
Mean ratings from Stage 2 (post-discussion), n = 4 participants per stakeholder group. Hover any marker for the exact mean and standard deviation, or click a marker or feature name to read its full description.
Did the panel converge? Stage 1 → Stage 2
Comparing the full Stage 1 panel (n=21) against the full Stage 2 panel (n=16) for every feature shows whether the discussion phase narrowed disagreement. A feature is considered at consensus when its interquartile range (IQR) is ≤ 1 on the 5-point scale.
Consensus outcome buckets four raw statistical labels into the colour scale above: consensus covers features that reached, strengthened, or maintained IQR ≤ 1; converging narrowed but hasn't reached it yet; stable/diverged stayed above the threshold or widened; consensus lost had IQR ≤ 1 in Stage 1 but not in Stage 2. Kendall's W rose from 0.232 to 0.300 across the two stages, indicating stronger, though still moderate, agreement after discussion.