Data analysis at Magic Toybox exists to provide reliable, actionable intelligence to vendors in such a way that yields verifiable, noticeable, and gainful results from them in terms of traffic and revenue. We further aim to be a means to show these results. This is accomplished through the combination of psychometric/behavioral data gathered through use of our cell phone apps, as well as economic data that informs our demographic understanding.
There are three core components to the data analysis process:
The system operates as a controlled evidence pipeline:
User and vendor observations
↓
Normalized sessions and outcomes
↓
Behavioral and operational features
↓
Privacy-safe venue cohorts
↓
Eligible hosting-action candidates
↓
Policy-based ranking
↓
Vendor suggestion with explanation
↓
Observed results and experiment feedback
The consumer application records observable facts such as:
The vendor application records facts such as:
These are immutable observations. The mobile applications do not decide that a user has a particular personality or that a vendor should host a particular game.
The backend turns individual events into normalized records such as:
This sessionization step makes otherwise fragmented mobile events comparable. For example, five objective events, a pause, three retries, and a terminal completion event become one coherent session with an understood duration and outcome.
The raw events remain immutable, while normalized facts can be deleted and rebuilt if processing rules change.
The first models calculate four bounded behavioral propensities:
Each feature includes:
These are descriptions of observed behavior under particular conditions—not permanent personality classifications. Low-confidence evidence is not treated like a well-supported result.
The initial vendor features measure:
These help the recommendation engine distinguish, for example, between a vendor that regularly operates multiple campaigns and one that has little experience configuring or activating them.
Individual user feature values are not shown to vendors. Instead, the system creates aggregate views scoped to the vendor’s own venues and campaigns.
Approved cohort dimensions can include:
This allows questions such as:
Vendor results must ordinarily represent at least 25 distinct users. Small cells are suppressed, related cells are suppressed when needed to prevent reconstruction, and counts are rounded. No user IDs, event IDs, individual features, segment memberships, exact coordinates, or exact small counts are returned.
The decision engine does not generate unrestricted prose. It chooses among approved, versioned action candidates linked to canonical operational records.
A hosting-action catalog can include actions such as:
The candidate record refers to the actual game, venue, campaign, offer, or catalog item. The recommendation system does not become the source of truth for those operational objects.
Before ranking, each candidate is evaluated as:
eligible;ineligible; orinsufficient_evidence.Hard rules can exclude an action because:
The explicit insufficient_evidence state is important. The engine should say that it does not yet know rather than interpreting missing information as poor performance.
A versioned policy ranks the remaining candidates using a controlled scoring structure:
final score =
base action score
+ approved feature contributions
+ segment or cohort-fit boosts
+ venue and time-context boosts
+ bounded deterministic exploration
The rule language supports:
For vendor-facing hosting suggestions, the evidence can combine:
Audience evidence
What privacy-safe user cohorts select, complete, repeat, redeem, or abandon.
Venue context
Place, zone, time, operating conditions, available products, and campaign goals.
Vendor behavior
The vendor’s campaign activity, game enablement history, and willingness to test audience configurations.
Operational constraints
Game availability, capacity, reward inventory, schedule, and venue suitability.
Historical outcomes
Results from prior recommendations and experiments.
The engine persists the contribution of each piece of evidence internally, so the decision can later be audited.
The vendor does not receive the internal score, individual user features, or individual segment memberships.
The response contains:
The vendor interface can translate safe explanation codes into statements such as:
Internally, the system retains a more detailed audit explanation, including rule outcomes, feature references, contributions, and reasons other candidates were rejected. Public responses deliberately omit scores, confidence values, internal failures, and individual segment states.
The vendor application records whether a suggestion was:
If the vendor implements it, subsequent user behavior provides the actual outcome:
These outcomes are attributed to the original decision item and rolled up idempotently. The next feature refresh and recommendation cycle therefore incorporates what happened after the suggestion.
Where evidence does not establish which configuration is better, the system can recommend or internally authorize a controlled experiment.
For example:
Variant A: cooperative 30-minute game, 5–8 p.m.
Variant B: cooperative 45-minute game, 5–8 p.m.
Assignment is server-controlled, deterministic, and sticky. Exposure is recorded only when the corresponding decision is actually delivered. Results can compare completion, selection, redemption, duration, or retention, subject to minimum samples and guardrails. A human must approve production policy changes; the system does not automatically promote a winner.
Suppose the system observes that, at a resort:
The engine could evaluate candidates such as:
After checking venue capability, hours, game support, campaign status, reward inventory, cohort size, and evidence confidence, it might return:
Recommended: Host the cooperative exploration configuration from 5–9 p.m., using a waterfront route and dining reward.
Why: Strong evening group participation, favorable completion within the proposed duration, and an opportunity to improve discovery of an underused venue area.
Evidence status: Sufficient for a bounded pilot; test against the current configuration before broader deployment.
The vendor sees the business recommendation and aggregate rationale. It does not see which individuals contributed to the conclusion.