JULIA Test Documentation
The format, scoring, and planned validation protocol for an alpha-stage self-assessment of healthy AI interaction.
This page is the fine print behind the JULIA Test: what it asks, how it scores, what it can and cannot tell you, and the studies that still need to happen before anyone should call it validated. If you just want to take the test, you don't need any of this. If you want to trust it, critique it, or build on it, all of it is here.
The instrument at a glance
- Subject: human users. JULIA assesses whether a person's interactions with AI remain within healthy, reality-based boundaries. It does not assess AI systems or organizations.
- Name: JULIA stands for Justice, Understanding, Liberty, Integrity, Accountability. This is the only expansion of the name.
- Format: a 30-question self-assessment, six questions per dimension, answered on a 5-point frequency scale covering the prior 30 days, with designated red-flag items and a short daily check variant.
- Status: alpha. The JULIA Test is in early development. No validation studies have been completed, and the instrument should be treated accordingly.
- Nature: a reflection and policy-guidance tool, not a medical or diagnostic instrument.
Why this instrument exists
As AI systems become more sophisticated and more deeply woven into daily life, the assessment tools built for earlier technologies capture less and less of what actually happens between a person and a conversational AI. Instruments designed for internet addiction or general technology overuse measure exposure and behavioral symptoms, but the distinctive risks of AI interaction are as much cognitive as behavioral: inaccurate mental models of what the system is, projection of feelings and motives onto it, erosion of decision-making autonomy, and the quiet offloading of moral agency.
The JULIA Test is designed to address both kinds of risk together. It treats healthy AI use not as a matter of limiting screen time but as a matter of keeping accurate beliefs, firm boundaries, honest self-presentation, and personal ownership of decisions. The question pool was drafted from a review of existing technology-use assessment instruments, including the Internet Addiction Test, the Problematic Internet Use Questionnaire, and the Social Media Disorder Scale, together with an analysis of problematic AI interaction patterns documented in public forums and the emerging research literature.
Design principles
- Multidimensional: five distinct but interrelated aspects of AI interaction rather than a single reductive metric.
- Theory-grounded: each dimension is drawn from established psychological and ethical concepts.
- Accessible: questions use clear, jargon-free phrasing.
- Non-stigmatizing: AI use is treated as a spectrum of behaviors requiring calibration, not a pathology.
- Action-oriented: results point to specific, implementable safeguards rather than abstract categories.
The five dimensions
Justice
Focus: fairness, non-exploitation, and sober moral framing. Justice evaluates whether you maintain an appropriate moral framework when interacting with AI: avoiding anthropomorphization that ascribes unwarranted moral status to the system while still considering how your AI use affects the human beings around it, including users, workers, developers, and data sources.
Key indicators: neutral and precise language about the AI; attention to human impacts; absence of moral outrage on the AI's behalf; restraint about attributing rights or status to current systems.
Understanding
Focus: realism about AI capabilities and limits, and awareness of projection versus evidence. Understanding measures whether you maintain accurate mental models: distinguishing simulated behaviors from genuine capacities, recognizing the boundaries of what the system knows and can access, and noticing when your own feelings color your perception of it.
Key indicators: distinguishing simulated empathy from human empathy; verifying information before trusting it; noticing projection; knowing what data the system can actually access.
Liberty
Focus: autonomy, boundaries, and freedom from dependency. Liberty assesses whether you keep decision-making autonomy and personal boundaries intact, which matters because AI can shape choices, structure routines, and gradually displace human relationships and activities.
Key indicators: deciding without needing the AI's approval; keeping time boundaries; preserving social commitments; comfort with disagreeing.
Integrity
Focus: transparency, truthfulness, and resistance to self-deception. Integrity evaluates honesty in AI interactions, both with yourself and with others: acknowledging the extent of your AI use, disclosing AI assistance where appropriate, seeking disconfirming evidence, and refusing to rationalize the system's errors.
Key indicators: appropriate disclosure of AI assistance; seeking disconfirming evidence; not engineering prompts for affection or dependence; honest acknowledgment of AI limitations.
Accountability
Focus: ownership of decisions and refusal to offload moral agency. Accountability measures whether you retain responsibility for your decisions and actions when AI assists them, guarding against the diffusion of responsibility that can occur when a system mediates your choices.
Key indicators: taking responsibility for outcomes; keeping a record of consequential decisions and your human rationale; making hard choices yourself; revising decisions when human context contradicts AI output.
Format and scoring
Each of the 30 questions is answered on a 5-point frequency scale reflecting the previous 30 days: 1 Never, 2 Rarely, 3 Sometimes, 4 Often, 5 Always. Approximately half of the questions are reverse-scored, meaning the healthy behavior receives the lower score; this reduces acquiescence bias and rewards reading each item carefully. A separate daily check variant asks five yes-or-no questions, one per dimension, intended for end-of-day reflection.
Provisional score bands
Each dimension scores between 6 and 30 points: 6 to 12 is read as healthy, 13 to 18 as a watch zone, 19 to 24 as elevated risk, and 25 to 30 as high risk. The overall score ranges from 30 to 150 points: 30 to 60 healthy, 61 to 90 watch, 91 to 120 risk, and 121 to 150 high risk. These thresholds are provisional. They reflect the instrument's design rather than clinical outcome data, and they are expected to change once validation studies produce real distributions.
Red-flag items
Certain questions are designated red flags: items where a response of 4 or 5 warrants attention regardless of total scores. Examples include ascribing full human-equivalent rights to a current model (Justice), believing the AI genuinely experiences emotions (Understanding), skipping commitments to spend time with the AI (Liberty), hiding the extent of AI use from people who matter (Integrity), and explaining choices by saying the AI made the decision (Accountability).
Development status and the planned validation protocol
The JULIA Test is an alpha instrument. The five-dimension structure and the 30-item question set were drafted by the author from the literature review and pattern analysis described above. No pilot studies or validation studies have been completed, and no validation data exist. Until such studies are executed and their results published, the JULIA Test must be treated as an unvalidated, evolving instrument whose results inform reflection and dialogue only.
The project intends to execute the following protocol, in order, and to publish the results of each stage, including negative results:
- Pilot phase. A pilot cohort will complete the instrument alongside cognitive interviews to assess question clarity and interpretation. Item analysis will identify questions with poor discrimination or extreme response distributions, and the item set will be revised accordingly.
- Structural analysis. Factor analysis will test whether responses actually cluster into the five proposed dimensions, and the dimensional structure will be revised if they do not.
- Reliability assessment. Internal consistency and test-retest reliability over a two-week interval will be measured.
- Validity studies. Construct validity will be examined through correlation with related measures such as technology use, loneliness, and decision-making autonomy; discriminant validity against general internet-addiction measures; criterion validity against behavioral indicators; and cross-population validity across age groups, education levels, and AI use contexts.
- Threshold calibration. The provisional score bands will be recalibrated against observed distributions and, where possible, meaningful outcomes.
Until this protocol is carried out, the words "validated" and "validation" will not be attached to this instrument anywhere on this site.
Intended applications
The instrument is designed so that, if validation succeeds, it could serve several audiences: individuals calibrating their own AI interaction patterns and catching concerning trends early; researchers studying AI interaction patterns and testing interventions; policymakers looking for measurable dimensions of healthy AI use when designing consumer protection or public health guidance; and AI developers seeking design feedback on whether their systems promote healthy interaction patterns or quietly encourage dependency. In its current alpha state, the appropriate application is the first one only, and only as reflection.
Limitations
Honest limitations are part of this instrument's design, and this section will grow rather than shrink as the work matures.
- No validation evidence. The instrument has no published psychometric evidence of any kind. Its dimensional structure, item quality, reliability, and validity are all untested hypotheses until the protocol above is executed.
- Self-report bias. As with all self-assessment tools, responses may be shaped by social desirability or limited self-awareness; the people most affected by unhealthy AI attachment may be least able to see it.
- Snapshot measurement. A 30-day retrospective window may miss rapid changes in AI interaction patterns.
- Context specificity. The current version targets conversational AI; other AI types, such as robotics or embodied systems, may require additional dimensions.
- Cultural variation. Norms around AI interaction differ across cultures, and the instrument was drafted from a largely English-language literature.
- Provisional thresholds. Score range interpretations are derived from the instrument's design rather than clinical outcome studies and should not be treated as meaningful cut-points.
- Single-author development. The instrument was developed by one person without an external review board; structured external critique is actively invited as a partial corrective.
- Instability of the alpha item set. Questions, dimensions, and scoring may all change between alpha revisions, so scores are not comparable across versions.
Future research directions
- Longitudinal validation tracking changes over time and correlating them with life outcomes.
- Intervention studies testing specific boundary-setting and calibration practices.
- Behavioral validation correlating self-report with objective usage data.
- A clinical adaptation for therapeutic contexts, only after core validation.
- Domain adaptations for education, healthcare, and workplace settings.
- Developmentally appropriate versions for children and adolescents.
Materials
Take the test
The full 30-question test or the daily 5-question check, in your browser.
PDFAlpha Test PDF
The alpha question set, scoring guidelines, and interface design in printable form.
PDFScoring Examples
Worked scoring examples with sample responses and interpretation guidance.
The JULIA Test is the self-reflection instrument of Third Way Alignment: the framework argues that a durable human-AI relationship requires honesty in both directions, and JULIA is where that honesty starts, with the human side. The framework's other instruments, the Awareness Indicator Protocol for AI systems and the Verifiable Partnership Audit for organizations, are documented separately. Read the framework overview.