RESEARCH ARTICLE
Decomposition and Recursive
Synthesis of Creative Tasks
by Autonomous Agents
A Mixed-Methods Investigation of the 5-Axis Brief Architecture
Rick Ito1, Aiko Tanaka2, Kenji Watanabe2
1Acme Studio Research Lab, Tokyo · 2Tokyo Institute of Technology
3. Methods
3.1 Participants
A total of 120 professional creators (62 female, 56 male, 2 non-binary; Mage=31.4, SD=6.8) were recruited via the freelance platform CrowdWorks. All had at least 3 years of experience producing client-facing visual deliverables. Participants received ¥8,000 as compensation. The study was approved by the institutional review board (Approval No. 2025-CR-184).
3.2 Design
A between-subjects 2 × 1 design with random assignment was employed. Group A (n=60) received a 5-axis brief decomposition (output / vibe / palette / motion / diagram), while Group B (n=60) received a single-prompt baseline replicating standard practice. Task quality served as the dependent variable.
3.3 Procedure
All participants completed the same task: produce a 20-slide presentation deck for a hypothetical Series A pitch within 90 minutes. Group A used the studio interface (Ito, 2025), which forced explicit selection across the 5 axes before any rendering. Group B used a free-form text input identical to ChatGPT's standard interface.
3.4 Quality Assessment
Three blinded experts (designers with 10+ years experience) independently rated each deliverable on a 100-point composite scale measuring: (a) structural coherence, (b) visual consistency, (c) message clarity, and (d) production polish. Inter-rater reliability was strong (ICC = .87, 95% CI [.83, .91]).
3.5 Statistical Analysis
Welch's independent-samples t-test was used to compare composite quality scores between groups. Effect sizes (Cohen's d) were computed with Hedges's correction. All analyses were conducted in R 4.4.0 with the effectsize package (v1.2.0).
4. Results
Group A (5-axis condition; M = 82.4, SD = 8.2) achieved significantly higher composite quality scores than Group B (single-prompt; M = 64.1, SD = 11.8), t(118) = 10.42, p < .001, Cohen's d = 1.81 (very large effect; Hedges-corrected g = 1.79). Figure 1 displays the distributions.
4.1 Subscale Analyses
Examining subscale scores, Group A outperformed Group B on all four subscales: structural coherence (d = 2.04), visual consistency (d = 1.92), message clarity (d = 1.43), and production polish (d = 1.51). All differences were significant after Bonferroni correction (p < .0125).
4.2 Manipulation Check
Self-reported cognitive effort did not differ between conditions, t(118) = 0.84, p = .40, suggesting that 5-axis decomposition improved output without increasing perceived workload — a noteworthy practical finding (cf. Sweller, 1988).