incomplete
The article below describes the experimental design component to Magic Toybox Games. Our games have their creative design, and experimental design components. The latter is how we take the data inputs and use them to develop and test hypotheses and thereby learn about our users.
Design & Operations Document
The Analytics & Survey Engine is responsible for:
- Designing and running field experiments tied to Magic Toybox games
- Generating and evolving hypotheses about player amusement and engagement
- Collecting survey responses from real-world play
- Scoring hypotheses based on outcomes
- Mutating and refining future hypotheses algorithmically
The system supports:
- Manual survey entry
- CSV ingestion from field teams
- Incremental learning via bandit-style optimization
- Separation of content, experiments, and analytics logic
This document is intended for:
- Engineers onboarding to the system
- Field team leads
- Data science / ML iteration planning
- Operations and governance review
| Term |
Meaning |
| Game |
A Magic Toybox experience (Treasure Hunt, Echo Garden, etc.) |
| Site |
Physical location where play occurs |
| Cluster |
Context bucket (game + site + group) |
| Hypothesis |
A vectorized assumption about amusement |
| Experiment |
A single field run testing one hypothesis |
| Survey Response |
Observed outcome from play |
| Reward |
Numeric score assigned to hypothesis |
flowchart LR
FieldPlay --> SurveyInput
SurveyInput --> Experiments
Experiments --> Scoring
Scoring --> BanditEngine
BanditEngine --> Hypotheses
Hypotheses --> Experiments
erDiagram
GAMES ||--o{ CONTEXT_CLUSTERS : defines
SITES ||--o{ CONTEXT_CLUSTERS : defines
CONTEXT_CLUSTERS ||--o{ HYPOTHESES : contains
HYPOTHESES ||--o{ EXPERIMENTS : tested_by
EXPERIMENTS ||--o{ SURVEY_RESPONSES : collects
QUESTIONS ||--o{ SURVEY_RESPONSES : answered_by
ANSWERS ||--o{ SURVEY_RESPONSES : selected
- question_bank
- answer_bank
These store canonical content:
- Questions are reused across surveys
- Answers are reusable and normalized
- Surveys reference subsets
This allows:
- Longitudinal comparison
- Cross-game analytics
- Future ML feature extraction
sequenceDiagram
participant API
participant Bandit
participant DB
API->>Bandit: Request hypothesis
Bandit->>DB: Select best hypothesis
Bandit-->>API: hypothesis_id
API->>DB: Create experiment
Each experiment records:
- Cluster
- Hypothesis
- Start time
- Status (open / closed)
Surveys are collected via:
- Manual UI entry
- CSV ingestion from field teams
sequenceDiagram
participant Field
participant API
participant DB
Field->>API: Upload responses
API->>DB: Store survey_responses
Rewards are calculated from survey responses:
- Likert scale normalization
- Weighted question importance
- Response completeness
- Optional outlier handling
flowchart TD
Responses --> Normalize
Normalize --> Weight
Weight --> Aggregate
Aggregate --> RewardScore
Rewards are stored on:
- Experiment
- Hypothesis (rolling average)
¶ 8. Bandit + Mutation Engine
The bandit engine balances:
- Exploitation: use known good hypotheses
- Exploration: try new variations
This avoids:
- Overfitting early assumptions
- Manual hypothesis tuning
¶ 8.2 Bandit State
Stored in bandit_state:
| Field |
Purpose |
| selection_count |
Times chosen |
| cumulative_reward |
Total score |
| last_selected_at |
Recency bias |
flowchart LR
Hypotheses --> ScoreEstimate
ScoreEstimate --> SelectBest
SelectBest --> Experiment
Selection strategies supported:
- ε-greedy
- Softmax
- UCB (future)
Triggered when:
- Selection count threshold reached
- Score plateau detected
- Exploration budget available
flowchart TD
ParentHypothesis --> MutateVector
MutateVector --> ChildHypothesis
ChildHypothesis --> InitializeBanditState
Mutation includes:
- Small random vector deltas
- Controlled variance
- Inheritance of metadata
Each hypothesis contains:
Vectors are stored using:
flowchart LR
Hypothesis --> Experiment
Experiment --> Survey
Survey --> Reward
Reward --> UpdateBandit
UpdateBandit --> Hypothesis
This loop runs continuously and improves:
- Game tuning
- Site matching
- Group targeting
- Running assigned experiments
- Collecting surveys post-play
- Uploading CSVs or entering responses
- Reporting anomalies
They are not responsible for:
- Hypothesis generation
- Scoring logic
- Analytics tuning
| Risk |
Mitigation |
| Bad early data |
Exploration bias |
| Small sample sizes |
Minimum N thresholds |
| Survey fatigue |
Rotating questions |
| Field variance |
Cluster isolation |
Planned extensions:
- Bayesian bandits
- Contextual bandits
- Offline batch retraining
- Multi-objective optimization
- Vendor-specific amusement models
All decisions are:
- Logged
- Reproducible
- Traceable to data
This supports:
- Academic rigor
- Partner transparency
- Long-term analytics trust
This engine provides:
- Structured experimentation
- Continuous learning
- Field-operable workflows
- Future-ready ML integration
It is intentionally modular, auditable, and human-in-the-loop.
- Lock scoring weights v1
- Define vector semantics formally
- Evaluate contextual bandits
- Design visualization dashboards
If you want, next I can:
- Convert this into a field-only runbook
- Add ML algorithm comparison appendix
- Generate architecture diagrams per module
- Create a README-lite version for vendors