These three sit together because they all point the same way—making models more useful with less human babysitting and fewer data leaks—but each is missing key pieces.
Runway Solaris is the flashiest, and I'd discount it the most. It generates interactive UIs frame-by-frame with no frontend code: clicks and drags drive the next frame directly. Runway's own study says users prefer it 61% to 24% over Claude Opus 5-coded interfaces on instruction-following, and 71% to 21% on behavioral naturalness. Cost-wise they claim it's orders of magnitude cheaper than standard video diffusion models. Both reference points are off: the preference study pits subjective "naturalness" against code, and the cost comparison is against already-expensive video diffusion, not normal web rendering. More importantly, there's no public testing, no pricing, no API—just curated demo videos and a waitlist. Reddit devs did the math: traditional frontend code, once deployed, costs nothing per click; frame-by-frame generation burns GPU on every interaction, so costs scale linearly with usage. Runway's own post lists four unsolved limits: text rendering is unstable, the model hallucinates plausible-but-wrong UI states, visual style drifts in long sessions, and pure pixel output can't work with screen readers. Those four alone block general software production.
Google WikiSkill tackles a real pain point: model weights are frozen after deployment, so every hard-won debugging lesson vanishes when you close the chat window. Their system distills agent failure logs into structured skill manuals, with a gating mechanism that only merges updates if validation scores strictly beat the historical best. On Gemini-3.5-Flash, accuracy went from 49.5% to 68.1%. Two counterintuitive findings matter more: dumping raw debugging logs directly into the executing agent's context dropped scores from 63.7% to 60.9%—noise hurts attention allocation. And skills distilled from a small model (Qwen-3.5-4B) tanked a larger model's accuracy from 50.5% to 18.1%—the small model's workarounds actively constrained stronger reasoning. The paper is CC BY 4.0 but has no official code repo, only a third-party reproduction.
GitHub GHES 3.22 is the most concrete and the least complete. It lets enterprises self-host Copilot CLI inside air-gapped networks, with admins managing model endpoints centrally. But it's explicitly labeled a technical preview, many features are disabled, and future versions may change things. The real catch: data only truly stays on-prem if the model endpoint is also inside the air gap—enterprises need to run their own inference infrastructure, which isn't a flip-a-switch thing.