Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
KickBench - Can LLMs predict Bundesliga games?
See how LLMs with web search predict Bundesliga games. This talk rates their reasoning and search behavior against real results, showcasing model differences.
I let LLMs with a very generic prompt and a websearch tool predict bundesliga matches and then rate the predictions based on real game results.
For every prediction of every LLM I save the full trace, so that I can compare what different models search for or how different models reason about bundesliga games.
- NextNext.js is the full-stack React framework: it delivers high-performance web applications via hybrid rendering and powerful, Rust-based tooling.This is the React Framework for production: Next.js enables you to build full-stack web applications with zero configuration and maximum efficiency. It supports a hybrid rendering approach (Server-Side Rendering, Static Site Generation, and Incremental Static Regeneration) for optimal speed and SEO performance. Key features include React Server Components, Server Actions for running server code directly, and the App Router for advanced routing and nested layouts. Developed by Vercel, it leverages Rust-based tools like Turbopack and the Speedy Web Compiler for the fastest possible builds and a superior developer experience.
- Anthropic APIProgrammatic access to Anthropic's Claude models (Opus, Sonnet, Haiku) for complex reasoning, vision, and tool-use applications.The Anthropic API delivers programmatic access to the Claude model family (Opus, Sonnet, Haiku), enabling developers to integrate state-of-the-art AI into applications. Use the Messages API for conversational tasks, leveraging Claude 3.5 Sonnet for balanced performance or Claude 3 Opus for complex analysis. Key features include Tool Use (function calling), Vision capabilities for image analysis, and a large 200K token context window for extensive document processing. This API provides a powerful, reliable foundation for next-generation AI projects.
- OpenAI APIOpenAI API: Your direct gateway to cutting-edge AI models (GPT-4o, DALL-E 3, Whisper), enabling scalable, multimodal intelligence integration into any application.The OpenAI API provides authenticated, programmatic access to a powerful suite of generative AI models. Developers leverage REST endpoints and official libraries (Python, Node.js) to integrate capabilities like advanced text generation (GPT-4o), image creation (DALL-E 3), and speech-to-text transcription (Whisper). This platform is engineered for scale, supporting millions of daily requests for tasks from complex reasoning to real-time customer support agents, ensuring your application gets reliable, state-of-the-art intelligence.
- Mistral SDKThe official Python and JavaScript interface for deploying Mistral AI’s frontier models like Mistral Large and Codestral.Mistral SDK provides the primary gateway to the Mistral AI API, enabling developers to integrate high-performance LLMs with minimal boilerplate. It supports synchronous and asynchronous clients for tasks including chat completion, embeddings generation, and function calling. By utilizing the SDK, teams can toggle between models like Mistral 7B for speed or Mistral Large 2 for complex reasoning (123B parameters) using a standardized interface. The library handles authentication via API keys and manages streaming responses for low-latency applications.
- MistralFrontier AI models (LLMs) from Paris: delivering top-tier performance and efficiency through open-source innovation and optimized architecture.Mistral AI is the Paris-based frontier AI startup, founded in April 2023 by ex-Google DeepMind and Meta researchers (Arthur Mensch, Guillaume Lample, Timothée Lacroix). We challenge opaque 'big AI' with a mission to democratize advanced models: focusing on open-source, efficiency, and performance. Our technology, including the 123B parameter Mistral Large 2 and sparse Mixture of Experts (MoE) architecture, consistently delivers state-of-the-art results at significantly lower costs. We provide enterprise-grade solutions (Mistral AI Studio, Le Chat) for custom deployment, fine-tuning, and full data control. We are scaling fast: a $14 billion valuation confirms our position as a global leader in accessible, powerful generative AI.
Related talks
More from the community
Breaking the Chains - an organization-theoretical approach to creating LLM applications
Berlin
Explore an organization-theory-inspired approach to building intuitive LLM applications. This project simplifies prototyping mid-to-large scale LLM pipelines with…
Plug-in hybrid: deterministic solving engine to combine with LLMs
Bremen
This talk explores using deterministic solving engines with LLMs to manage complex rule sets, converting quantitative models into…
They Bet, They Lose, They Learn: The Agents That Rewrite Their Own Playbook
Cologne
An agent learns to predict World Cup 2026 game outcomes by using and iteratively improving its own playbook…
Utilizing Synthetic Datasets for Sales prospects
Bremen
Learn how LLMs generate synthetic B2B sales scenarios and tailored insights, enabling fast, practical preparation for sales calls…
YourBench - Benchmarking LLMs on your Data
Dublin
This talk demonstrates how to create a dataset and use YourBench to efficiently evaluate the performance of large…
Current trends + advances in Open LLMs - brief roundup by DiscoResearch
Munich
Overview of recent open‑LLM advances, demo of DiscoLM German, lessons from large‑scale fine‑tuning projects, and practical guidance for…
Compose Email
Loading recent emails...