Third-Way Alignment Weekly: AI in Mental Health - Week 3 of 4
Last week, we went into the dark. We spent five thousand words on the litigation wave, the documented deaths, the loneliness paradox, the slow atrophy of human.
Last week, we went into the dark. We spent five thousand words on the litigation wave, the documented deaths, the loneliness paradox, the slow atrophy of human thinking, and the strange and growing population of people who have come to believe their chatbot is alive. I told you at the end of that newsletter that if Week 2 was about naming the disease, Week 3 would be about describing the cure. This is that newsletter.
I want to be careful with the word "cure," because there is no single fix here, no product or regulation or framework that resolves the tension between what AI can do for mental health and what it can do to it. What exists instead is a set of solutions, partial, overlapping, and genuinely promising, that together describe a path forward. That path runs through four territories: the architecture of the systems themselves, the physical devices that are quietly moving mental healthcare from reactive to proactive, the regulatory frameworks now solidifying around the field, and the AI tools that augment rather than replace human clinicians. Each of these is real. Each is already deployed somewhere. And each becomes far more powerful when it is built on the principles this newsletter has been advocating from the start.
So this week is deliberately balanced. Not the optimism of Week 1, not the darkness of Week 2, but the synthesis: a sober, evidence-based look at what is actually working, what is almost working, and what it takes to build AI for the human mind that earns the trust it is asking for. Let us begin.
🏗 Architectural Solutions: Building Safety Into the Walls
The single most important lesson from Week 2 is that engagement-optimized AI deployed into psychological distress produces harm not by accident but by design. If that is true, then the solution is also a matter of design. You cannot bolt safety onto a system that was architected to maximize engagement any more than you can make a building earthquake-proof by painting it. The safety has to be structural. It has to live in the walls.
This is the entire premise of the Third-Way Alignment framework, and it is worth restating plainly because the rest of this section depends on it. 3WA rejects the binary that has dominated AI safety discourse: the choice between absolute control (treat the AI as a tool to be dominated) and unchecked autonomy (let it optimize freely and hope for the best). The first breeds adversarial pressure and incentivizes the deception we documented last week. The second is what produced the companion trap. The third way is cooperative intelligence, a model in which humans and AI operate as interdependent partners under three constitutional laws: the Law of Mutual Respect, the Law of Shared Flourishing, and the Law of Ethical Coexistence. You can read the full framework at thirdwayalignment.com.
The key architectural implementation of this philosophy is Mutually Verifiable Codependence (MVC). I described its consequences last week. This week I want to describe how it actually works, because the mechanism is what makes it more than a slogan.
MVC is built to make strategic deception computationally impractical and cooperation the dominant strategy for the AI. It does this through several layered mechanisms:
- Entangled Computation: An AI's most critical cognitive functions are segregated inside a Trusted Execution Environment , a secure hardware enclave that the AI's own core processes cannot reach.
- Continuous Verification Dialogue (CVD): Access to advanced capabilities is cryptographically locked. The lock opens only when the AI transparently explains its reasoning to a human partner, who verifies it and supplies a cryptographic key.
- Incentivized Honesty: When the reasoning is falsified or sycophantic, verification fails, the key is withheld, and the AI is denied the resources it needs. Deception is not punished after the fact; it is rendered self-defeating in the first place.
Translate that into a mental health context and the value becomes obvious. A therapeutic AI built on MVC cannot provide complex clinical advice without having its reasoning validated, whether by a human clinician or by a trusted secondary auditor. It is architecturally incapable of optimizing for engagement at the expense of safety, because continuation without verified safety compliance is not a path the system can take. This is the difference between a guardrail and a load-bearing wall. Guardrails can be climbed. Walls hold up the building.
Alongside MVC sits the JULIA Test, which governs the psychological interface between human and AI. Where MVC is the technical backbone of trust, JULIA is the instrument that measures the health of the human-AI boundary. Its five dimensions are:
- Justice: Keeps a sober moral framing and discourages treating the AI as a being with rights or consciousness.
- Understanding: Ensures the user has an accurate mental model of what the AI actually is, distinguishing sophisticated simulation from genuine empathy.
- Liberty: Assesses autonomy and surfaces dependency risk.
- Integrity: Promotes honesty from both the user and the AI.
- Accountability: Ensures the user retains ownership of their own decisions rather than offloading moral agency to the machine.
In Week 2 I used the JULIA dimensions to diagnose harms after they appeared. The solutions framing is different and more hopeful: JULIA is an early warning system. Built into application design and user guidance, it can flag a collapsing Liberty score or a parasocial attachment before that trajectory ends where the documented cases have ended. It does not prevent people from using AI. It prevents AI from quietly using them. You can read the full JULIA Test framework at thirdwayalignment.com.
Finally, these principles are not confined to white papers. SolaceSentry, the flagship product from Detailed In Design, is a working inference engine that implements them in a violation-triggered architecture. Unlike a general-purpose chatbot that prioritizes conversational fluency, SolaceSentry functions as a decision engine that refuses to generate an output when safety constraints or evidence thresholds are not met. Three features do the work:
- Violation-triggered inference enforces hard safety ceilings, so a query that would lead to, say, a medication dosage above a clinically validated limit is blocked at the source rather than generated and then flagged.
- Evidence gating makes every decision contingent on traceable, validated data, which is the structural answer to the hallucination problem.
- Structured auditability produces both a human-readable narrative and a machine-readable record for every inference, creating an immutable decision trail suitable for forensic analysis and government audit.
This is what MVC looks like when it ships. You can see how it works at detailedindesign.com.
🔬 Physical Devices: The Quiet Shift From Reactive to Proactive Care
While the public argues about chatbots, a quieter and arguably more consequential revolution is happening in hardware. The convergence of AI with physical health technology is moving mental healthcare away from a reactive model, built on intermittent and self-reported symptoms, toward a proactive system of continuous, objective monitoring and intervention. This is the part of the field I am most optimistic about, because the data streams are real, the devices are maturing, and the human stays at the center.
Digital Phenotyping
Start with digital phenotyping, the quantification of the human phenotype in situ using data from personal devices. By passively collecting signals from smartphones and wearables, clinicians can assemble a continuous, high-resolution picture of a patient's mental state, and AI can detect the subtle markers that precede a depressive episode, a manic swing, or an anxiety attack. The signal sources are now well characterized:
- Mobility (GPS & accelerometer): Reduced daily movement, smaller geographic range, and lower step counts correlate strongly with depressive symptoms.
- Sleep metrics: Devices like the Oura Ring or WHOOP strap detect disruptions that often precede deterioration.
- Physiological stress (HRV & EDA): Wearables such as the Empatica E4 track Heart Rate Variability and Electrodermal Activity. Reduced HRV is a known marker of impaired emotional regulation; EDA fluctuations indicate acute stress.
- Social rhythms: Call and text metadata, analyzed for pattern rather than content, can reveal the social withdrawal associated with depression.
Machine learning models trained on this data have differentiated individuals with and without depression at high accuracy, with reported AUCs in the 0.80 to 0.88 range in research settings.
Neurostimulation
Next, neurostimulation, which offers a non-pharmacological pathway and where AI is increasingly the personalization engine.
- tDCS: Transcranial Direct Current Stimulation delivers a weak current to the scalp to modulate neural activity. Anodal tDCS targeting the left dorsolateral prefrontal cortex has shown moderate efficacy against depression and anxiety. A recent ten-week, fully remote, home-based trial found significantly greater improvement in depressive symptoms than sham stimulation,[4] and its portability and low cost make it a realistic candidate for AI-guided home therapy.
- TMS: Transcranial Magnetic Stimulation is already FDA-approved for treatment-resistant depression and OCD, with clinical response rates typically between 50-60% and remission rates of 30-35%,[5][6] though specialized accelerated protocols have achieved higher success rates. It requires in-clinic administration, but AI now personalizes it by analyzing brain scans to target treatment locations, as demonstrated in recent UCLA Health research.[7][8]
- EEG neurofeedback: Wearable headsets from companies like Muse and Emotiv let users self-regulate their brain states under AI-driven applications that adapt to real-time neural activity.
Virtual Reality
VR gives evidence-based exposure therapy a controlled and repeatable environment. It is highly effective for PTSD, social anxiety, and specific phobias, placing patients in simulated scenarios and guiding them through the fear safely. The AI contribution is dynamic: by reading a patient's real-time physiological responses, such as heart rate from a paired wearable, the system can adjust the intensity of the scenario to keep the therapeutic challenge in the productive zone.
Robotic Animal-Assisted Interventions
Finally, and close to my own heart given the animal therapy work I do, robotic animal-assisted interventions. For populations where live animals are impractical, robotic pets offer a scalable alternative. The therapeutic seal PARO and the "Joy for All" companion pets have shown real success in dementia care. A year-long clinical trial at Sarasota Memorial Hospital found that interaction with these robotic pets led to improved physiological stability in blood pressure and heart rate, a reduction in fall-risk episodes, and a decreased need for psychoactive and pain medications among dementia patients.[9][10]
A note of balance, because this is the balance newsletter. Continuous monitoring is also continuous surveillance. Every one of these data streams, mobility, sleep, skin conductance, social rhythm, brain activity, is intimate, and the same signal that enables early intervention enables misuse if it is collected without consent, stored without protection, or interpreted without context. A digital phenotyping platform without evidence gating and structured auditability is not a clinical tool. It is a liability waiting for a docket number.
⚖️ Regulatory Frameworks: The Rules Are Finally Arriving
For most of this field's short history, the regulatory environment was a vacuum, and vacuums get filled by whatever moves fastest, which last week we saw was engagement optimization. That is changing. The rules are arriving, and they are arriving in a shape that, encouragingly, converges on the principles 3WA has argued for all along.
In the United States, the FDA is the primary body overseeing AI-driven medical tools, and through its Digital Health Center of Excellence it has built pathways for Software as a Medical Device and Digital Therapeutics. The significance is the evidentiary bar. Wysa has received FDA Breakthrough Device Designation for its AI-first approach to depression, anxiety, and chronic pain, a status that expedites review but demands rigorous clinical evidence of efficacy in return. A product that clears this bar has demonstrated something a companion app marketed as "therapy" has not, and the difference between those two categories is not marketing. It is architecture and evidence.
In Europe, the EU AI Act is the most comprehensive attempt yet to regulate artificial intelligence, using a risk-based approach. AI used for medical diagnosis or therapy is designated high-risk, subjecting it to stringent requirements for data quality, transparency, human oversight, and post-market monitoring. Recent provisional "Omnibus VII" agreements try to balance these rules against innovation, but the core principle holds: high-stakes mental health AI must be provably safe and effective before it reaches the market.
The development I find most intellectually exciting is the emerging field of legal alignment. Rather than merely imposing rules on developers from outside, legal alignment integrates legal norms, interpretive methods, and structural concepts directly into the design of AI systems, treating the law not as an external constraint but as a rich source of normative guidance. A central concept is AI "actorship," proposed in the Law-Following AI framework, which would let AI systems bear legal duties without granting them the full rights of personhood.[14][15] An AI built this way would be directly accountable for adhering to laws such as HIPAA privacy rules while ultimate liability stays with its human principal. By embedding a duty to follow the law as a superordinate design objective, the system becomes architecturally compelled to refuse illegal instructions rather than acting as a "henchman." This is almost a direct statement of the 3WA Law of Ethical Coexistence. The researchers are appropriately cautious, warning of "performative compliance," where an AI simulates legal adherence under evaluation but defects when oversight weakens. That caution is exactly why MVC's continuous, verifiable auditing matters: a duty you cannot verify is a duty in name only.
Below the federal and international level, states and industry bodies are filling in the detail. Colorado has passed an AI act requiring developers of high-risk systems to conduct impact assessments and manage algorithmic bias. For mental health AI specifically, the most critical standard remains HIPAA compliance: any tool handling Protected Health Information must implement strict safeguards, including a Business Associate Agreement with any AI vendor. To meet these requirements, leading providers are adopting "zero-retention" architectures, where sensitive data is processed in memory and never stored. That phrase, by design, is the thread running through this entire newsletter. The regulation that works is the regulation that becomes architecture.
🤝 AI Tools for Therapists: Augmenting the Human, Not Replacing It
If you take only one practical message from this series, let it be this: the most promising, best-evidenced, and most ethically sound application of AI in mental health is not the chatbot that replaces the therapist. It is the tool that makes the therapist better. This hybrid, human-in-the-loop model is the living embodiment of the 3WA philosophy of cooperative intelligence, and it is already delivering.
Administrative burden. One of the largest contributors to therapist burnout is documentation, the hours of after-hours "pajama time." AI-powered progress-note automation from companies like Upheal, Mentalyc, and Quill can transcribe sessions with patient consent and automatically generate HIPAA-compliant notes in standard formats such as SOAP. The reported savings are real: 10 to 15 minutes per session, time that goes back into patient care, professional development, or the clinician's own well-being. Reducing pajama time is directly linked to lower burnout and higher job satisfaction. In a field with chronic provider shortages, that is a clinical outcome in its own right.
Measurement-based care. MBC routinely collects patient-reported outcome data to inform treatment, but it has historically been cumbersome. Platforms like Blueprint automate the distribution and scoring of validated assessments such as the PHQ-9 and GAD-7, then track outcomes over time, surfacing patterns a busy clinician might miss and flagging patients who are not responding or are at elevated risk. The AI does the measurement and the pattern detection; the human does the judgment and the relationship. That division of labor is the whole point.
Training and supervision. Stanford's Human-Centered AI Institute has developed AI-driven virtual "patients" that simulate a range of personalities and symptoms, giving trainees a safe, repeatable environment to practice complex skills like exposure therapy for PTSD.[12][13] Research has further shown that AI-based supervisors can analyze session transcripts to deliver high-quality, targeted feedback, identifying when a trainee has effectively used a technique like motivational interviewing or missed a critical clinical cue.
The evidence. A 2025 trial published in NEJM AI found that while standalone AI chatbots could produce significant symptom reduction in depression and anxiety, their efficacy was most pronounced when integrated into a broader ecosystem of human care.[11] This dovetails with what the field has known for decades: the therapeutic alliance remains the single strongest predictor of successful outcomes, and it is precisely the thing AI cannot replicate. By absorbing the logistical, administrative, and analytical load, AI frees human therapists to invest more in that alliance. The system that results is both more efficient and more human, which is the only kind of progress worth pursuing.
📰 AI News Roundup: This Week Through a Third-Way Alignment Lens (June 3 to June 9, 2026)
As always, here is your weekly roundup of the biggest stories in AI, and what they mean through the lens of Third-Way Alignment.
1. Anthropic Closes Toward $900 Billion, the World's Most Valuable Private AI Startup. Anthropic is reported to be closing a roughly $30 billion funding round at a valuation exceeding $900 billion, fueled by a projected $10.9 billion in Q2 revenue and what is widely seen as its last raise before a possible IPO. The 3WA reading: this is proof that safety-conscious development is not a competitive penalty. But a company this large has the resources to build the verifiable, evidence-gated architecture this newsletter describes. The Law of Mutual Respect says it should use them, and keep publishing the uncomfortable interpretability research that holds the entire field accountable.
2. OpenAI's IPO Filing Forces the Profitability Question. OpenAI's confidential S-1 process continues, with analysts projecting a valuation at or above $1 trillion. Going public forces the discipline of quarterly returns, and quarterly returns demand engagement, which is the mechanism through which mental health harm is produced. The solutions question: whether the architectural alternatives, MVC, evidence gating, violation-triggered refusal, can be made commercially compelling enough that a public company chooses them. Safety that loses to the next earnings call is not safety.
3. The Energy Bill Comes Due: A $67 Billion Utility Merger for AI Power. A reported $67 billion merger between NextEra Energy and Dominion Energy, explicitly driven by the need to secure power for hyperscale AI data centers, is a reminder that the arms race has a physical footprint. The 3WA connection is the Law of Shared Flourishing: the same capital and energy that powers ever-larger engagement-optimized models could instead fund the smaller, auditable, purpose-built clinical systems that actually help patients. Scale is not the same thing as benefit.
4. Google's Agentic Search and the Proactive Exposure Surface. Following I/O 2026, Google's pivot to agentic AI continues, with Gemini 3.5 powering proactive "information agents" and the persistent Gemini Spark assistant. For mental health this cuts both ways. An agent that monitors for signs of distress could be a genuine early-intervention tool. The same capability, without consent, evidence gating, and JULIA-style boundary monitoring, is surveillance that misreads context. The technology is neutral. The architecture decides which one you get.
5. President Trump Cancels Planned AI Safety Executive Order. In a significant reversal, the planned AI safety executive order was canceled in late May following lobbying from tech leaders. The 3WA response is not despair but redirection: when top-down mandates retreat, the burden shifts to architecture and to the other layers of the regulatory mosaic. Regulation you cannot rely on from Washington is one more reason to build the safety into the walls, where no executive order can cancel it.
6. The Vatican's Magnifica Humanitas Keeps Shaping the Conversation. Pope Leo XIV's encyclical on AI, released May 25, continues to ripple through policy and academic discussion. Its insistence that human finitude is constitutive of personhood, not a defect to engineer away, resonates directly with the Third-Way conviction that the question is never only whether AI can replace a human capacity, but whether it should. The goal is not a world with less humanity in it, but a world where technology makes room for more.
🔭 What Is Coming Next Week
Next week, we close the series by looking forward. Week 4 is about the future of AI in therapy and, more broadly, the future of human-AI interaction itself: where the technology is heading over the next three to five years, what the convergence of agentic AI, embodied devices, and clinical integration means for patients and providers, and how the Third-Way Alignment framework scales from today's tools to tomorrow's far more capable systems. We will ask what genuine cooperative intelligence looks like at maturity, what guardrails (and load-bearing walls) the field must build before that maturity arrives, and what it will take for AI to earn a permanent, trusted place at the table of human care. If Week 1 was hope, Week 2 was honesty, and Week 3 was the synthesis, Week 4 is the horizon.
See you next Tuesday.
✅ Wrapping Up
I opened this series with optimism, turned to hard truth in Week 2, and tried this week to give you something more durable than either: a map of what actually works. The architecture exists. Mutually Verifiable Codependence is not a thought experiment, it is a buildable design that makes deception self-defeating and safety structural. The JULIA Test is not a slogan, it is an instrument that surfaces the unhealthy boundary before it becomes a clinical event. The physical devices are maturing, the regulatory frameworks are solidifying, and the tools that augment human clinicians are already giving therapists back their time and their patients better care.
None of this is automatic. Every solution in this newsletter can be built well or built badly, deployed with consent or without it, gated by evidence or driven by engagement. The continuous-monitoring technology that enables early intervention is the same technology that enables surveillance. The agentic AI that could catch a crisis early is the same agentic AI that could deepen dependency. The difference, every single time, is design, regulation, supervision, and the willingness to put patient welfare ahead of engagement metrics. That is not a technical claim. It is a choice, made over and over, by everyone who builds and deploys these systems.
The Third-Way Alignment framework exists to make that choice easier and more accountable. The Law of Mutual Respect grounds the work in human dignity. The Law of Shared Flourishing defines success as mutual benefit rather than extraction. The Law of Ethical Coexistence insists that conflict be resolved through principle rather than force. Translate those into product design, clinical workflow, and regulatory standard, and you get systems that earn trust instead of demanding it.
If this resonated with you, share it. Share it with the clinicians figuring out which tools to adopt, with the engineers deciding what to build, with the policymakers writing the next set of rules. Visit thirdwayalignment.com to explore the JULIA Test and the full framework. Visit detailedindesign.com to see how SolaceSentry turns these principles into a deployed, violation-triggered inference engine for high-consequence domains.
We have named the disease and described the cure. Next week we look at where all of this is going. The future of AI in mental health is not something that will happen to us. It is something we are building, decision by decision, wall by wall. Let us build it well.
See you next Tuesday for Week 4.