Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Getting UI Agents to run reliably in enterprise environments
Learn how to run UI agents securely in enterprise sandboxes with policy enforcement and efficient trace analysis, demonstrated through a claims processing workflow.
We built a system that enables you to run agents in secure sandboxes with security policies and easily monitor/evaluate their progress. The demo shows this along the path of work inside a claims processing task.
Systematically measures AI performance, token usage, and latency using custom semantic evaluation criteria.
- ClaudeClaude is Anthropic's flagship family of large language models (LLMs): a high-performance, Constitutional AI system built for safety, complex reasoning, and expert-level collaboration.Claude is a next-generation AI assistant developed by Anthropic, a research firm prioritizing AI safety. The models (including Opus, Sonnet, and Haiku) leverage Constitutional AI to ensure helpful, honest, and harmless outputs, a key differentiator from competitors. Claude excels at complex enterprise tasks: processing massive context windows for in-depth data analysis, generating and reviewing code, and providing expert-level summarization for documents up to 200,000 tokens. It is deployed as a conversational chatbot and via API, offering scalable AI solutions for developers and businesses.
- OpenShellOpenShell is the definitive open-source successor to Classic Shell, restoring the functional Start menu and Windows Explorer toolbars to Windows 10 and 11.OpenShell delivers a high-performance replacement for the modern Windows UI, focusing on productivity through granular customization. It reinstates the classic two-column menu layout, provides skinning support for taskbars, and adds essential file management shortcuts to Windows Explorer. By utilizing a lightweight C++ codebase, it maintains a minimal memory footprint while offering power-user features like custom start buttons and advanced search filtering. It is the industry standard for users requiring the reliability of the Windows 7 interface on NT 10.0+ systems.
- PlaywrightPlaywright is the Microsoft-developed, cross-browser automation framework: it drives Chromium, Firefox, and WebKit with one unified API for fast, reliable end-to-end testing.Playwright delivers robust, cross-platform end-to-end testing, supporting all major rendering engines: Chromium, Firefox, and WebKit. Launched by Microsoft in January 2020, its core strength is a single API for multiple languages (TypeScript, Python, Java, .NET). The framework eliminates flaky tests through automatic waiting and provides full test isolation by creating a new browser context (a brand-new browser profile) for each test. Key tooling includes Codegen for recording actions and the Trace Viewer for deep post-mortem analysis of test failures (screencasts, live DOM snapshots). This architecture ensures reliable, high-speed execution across Windows, Linux, and macOS.
- ElluminateA synchronous virtual classroom engine acquired by Blackboard to provide low-bandwidth VoIP and interactive whiteboarding for global education.Elluminate Live! defined the early 2000s synchronous learning space before its 2010 acquisition by Blackboard for 116 million dollars. The platform optimized data transfer to ensure smooth audio and multi-user whiteboarding even on dial-up speeds. Major institutions (such as the Florida Virtual School) relied on its breakout rooms and session recording capabilities to manage thousands of remote students. While the original brand is now part of Blackboard Collaborate, its architectural focus on real-time pedagogical tools: polling, hand-raising, and desktop sharing: remains the industry benchmark.
Related talks
More from the community
The Quest for Tasks Frontier Models don't one-shot.
Berlin
Explore the challenges of creating complex evaluation environments for frontier models, discussing successes, failures, and pitfalls to avoid.
AI-Agents that learn from humans-in-the-loop
Nürnberg
Learn how AI agents proactively identify and acquire knowledge from employees, and delegate tasks, using Graph RAG for…
telli's internal agents infrastructure
Berlin
Learn how internal AI agents provide company-wide access for debugging, product information, bug research, and infrastructure monitoring, inspired…
Anita: A Proactive Synthetic Organism
Cologne
Explore Anita, a proactive AI architecture modeled on biological organisms. See how continuous stimuli trigger reactions and actions,…
Building Effective Agents
Brussels
Learn how to build effective agents using the Claude Agents SDK, with practical lessons learned from real-world use…
Running unattended agent jobs in production: scoping, isolation, and an overseer
Karlsruhe
Learn how to reliably run agent jobs in production using a system with scoping, isolation, and an overseer…
Compose Email
Loading recent emails...