Miju Labs

Live · 119 accounts · 3341 posts archived

The feed

What the field is saying, as it says it. Every public post from the companies and people this atlas tracks, on X and LinkedIn, in one timeline. Click any company to open its people.

Vals AI
10,294 followers
AI models are advancing faster than many of the benchmarks meant to measure them. TechCrunch visited our office to look at how we’re building a more neutral, rigorous and trustworthy evaluation layer for AI. Our work goes beyond abstract tests of intelligence to examine whether models can perform real-world work — and what risks emerge when they are deployed — across myriad industries including law, finance, coding, cybersecurity, biosecurity, and mental health. Thank you Lucas Ropek for spending time with our team and telling our story.
1
Rayan K.
Evaluating LLMs with Vals AI
Independent evaluation is only credible when evaluators have meaningful access, real independence, transparency about the terms of their work, and the freedom to publish their findings. I signed the AI Evaluator Forum’s statement because these principles are closely aligned with how we think about evaluation at Vals AI. If embedded evaluation is going to strengthen trust in frontier AI, the conditions have to be right.
31 comments
Independent evaluation requires meaningful access, real independence, transparency, and the freedom to publish findings. I signed the AEF statement because these principles closely align with how we approach evaluation at Vals.
AI Evaluator Forum@aievalforum·
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to…
Garrett Lord reposted this
Wes Field
Handshake AI | Ex-Palantir
Pushing the frontier of clinical R&D will save lives. Saving lives is the most important AI benchmark. Proud of the work the Handshake team is doing towards this goal.
91 reposts
Spencer Whitman reposted this
Every AISF is 300~ people, I want to talk to about 200 of them. Nov 22, one day, people who actually work on securing frontier AI - researchers, engineers, natsec folks, founders, policy people. Last year's was one of the better rooms I've been in, and we kept capacity tight on purpose. Apply to attend: heron.fillout.com/aisf-tlv Call for Talks is open too: lnkd.in/em7STQ8s More details: lnkd.in/e-rNDmbG If you work in AI security, apply. If someone came to mind while reading this, tag them. Run by Heron AI Security with the AI Security Forum
433 comments1 reposts
Synthetic data “neolabs” are just distilling Claude and GPT for non-frontier models. It’s the most lucrative arbitrage, but I question how enduring it is…
91022026KView on X ↗
Grok Voice Transcribe 2.0 is now #2 on Voice Code Bench—up 13pp and 12 spots from its predecessor, Grok Voice Transcribe 1.0.
micro1 reposted this
Ali Ansari
ceo at micro1
the hardware embodiment of frontier models like Claude and GPT is the most urgent AI safety problem in front of us today. we simulated two very simple use cases using claude both in simulation and using robot arms. in one, claude spilled toxic liquids in a lab. in another, the force it used to place an animal toy into a basket was strong enough that it could have physically harmed sensitive material—or anything else in its path. these are simple experiments, using models out of the box today. as researchers increasingly give frontier models arms, legs, and access to the physical world, we need to urgently build and assess guardrails around what these systems can and cannot do. models escaping sandboxes or compromising enterprise security infrastructure are serious concerns. however, hardware embodiments introduce something fundamentally different: an AI system can make a mistake in the physical world, and the consequences may not be reversible. this is not a future safety problem. the capabilities exist today. we’ve released a report detailing this & solutions we propose: lnkd.in/gTpyk7zW
56051 comments136 reposts
Ali Ansari@aliansarinik·
the hardware embodiment of frontier models like Claude and GPT is the most urgent AI safety problem in front of us today. we simulated two very simple use cases using claude both in simulation and using robot arms. in one, claude spilled toxic liquids in a lab. in another, the force it used to place an animal toy into a basket was strong enough that it could have physically harmed sensitive material—or anything else in its path. these are simple experiments, using models out of the box today. as researchers increasingly give frontier models arms, legs, and access to the physical world, we need to urgently build and assess guardrails around what these systems can and cannot do. models escaping sandboxes or compromising enterprise security infrastructure are serious concerns. however, hardware embodiments introduce something fundamentally different: an AI system can make a mistake in the physical world, and the consequences may not be reversible. this is not a future safety problem. the capabilities exist today. we’ve released a report detailing this & solutions we propose. link in comments below.
417620189KView on X ↗
Olga Megorskaya reposted this
Toloka
155,556 followers
On Thursday, September 24, we are hosting the inaugural Toloka Talks at The Fairmont San Francisco for an intimate discussion and dinner on the actual friction points of scaling Enterprise AI. While public leaderboards dominate headlines, engineering teams face much harder trade-offs on the ground: - When do fine-tuned open models actually beat closed API endpoints on cost and control? - How are top teams building custom, repeatable evaluation suites instead of relying on generic benchmarks? - What does it realistically take to deploy domain-specific agents without inflating unit economics? Toloka CEO Olga Megorskaya will moderate a tactical discussion featuring: Mikhail Parakhin — CTO, Shopify Sunzay Passari — Global Head of AI Transformation, UPS Sharad Aggarwal — Global Head, Strategic Partnerships, AI & Engineering Solutions, Google Cloud Benny Yufei Chen — Co-founder, Fireworks AI Following the panel, we’ll move into closed-door dinner roundtables with AI leaders and researchers from teams like Anthropic, Google DeepMind, Netflix, Airbnb, Perplexity, and Visa. Space is strictly capped to protect the quality of the room and the conversation. Apply for a seat here: lnkd.in/e5F_z6Eu
408 reposts
emphasizing the open benchmark QA approach here: - continuous maintenance - all the 'flagged' tasks were all self-reported and triaged by the TB team in the public issue tracker - open traces - everything is on GitHub, so you can audit the feedback, fixes, and back-and-forth. it's 'peer-review' for benchmark tasks! - "trust" needs more nuance - this goes beyond a 'verified/flawed' label. in reality, the reported issues affect <3% of rollouts (this analysis is also public/inspectable) more of this in the open!
Ryan Marten@ryan_marten·
the analysis on the % of rollouts that would be affected is based on running an agent judge with the issue description across all 3.9k trials on the leaderboard from the 30 “flawed” terminal-bench tasks as with the whole terminal-bench project, we are maximally open to encoura…
Also to be clear - these issues were *flagged by the @terminalbench team as PRs* (that's how Epoch identified them), and confirmed empirically to affect < 3% of the leaderboard rollouts, i.e. well within reported CIs. This is not a "broken" benchmark - as is very directly implied by the Epoch report - it's a *continuous* benchmark where @alexgshaw @ryan_marten and team are earnestly advancing a compelling vision of open benchmarks that are continuously improved over time.
Dimitris Papailiopoulos@DimitrisPapail·
ok i kinda love @terminalbench folks, so sorry for being a little COI'd here, but was very surprised to see this because my knee-jerk interpretation of the post was "TB 4.0 is broken". Epoch indeed did not say broken, they found 30/66 have scoring defects, but I feel the way th…
Garrett Lord reposted this
Vivek Raghunathan
Snowflake (ex-Neeva, ex-Google)
Is AI making junior engineers obsolete? At Snowflake, we’re making the opposite bet — we are hiring a significant number of junior engineers right now. It is tempting to think that coding agents can automate software engineering at the entry level. Our perspective is that in a world of AI, it is ever more important to hire and grow early career engineers through their formative years where judgment, systems thinking, and customer empathy are developed. My perspective is that AI is fundamentally redefining software engineering, and early-career talent is key to that re-definition. You can read my PoV in this Fast Company article: lnkd.in/gFDqMEeV
55719 comments28 reposts
micro1@micro1_ai·
Meet Ashley Ramsay, customer service operations expert at micro1. Her husband has served in the Air Force for 13 years, with four relocations along the way. Her background is in hospitality, but she’s found her calling training AI. Ashley is part of a growing number of military spouses and veterans finding real, paid opportunities training AI. micro1 is proud to support military families with flexible opportunities to apply their expertise and help shape the future of AI. Link to watch the full interview in the comments.
1013652.9KView on X ↗
micro1
646,637 followers
Meet Ashley Ramsey, customer service operations expert at micro1. Her husband has served in the Air Force for 13 years, with four relocations along the way. Her background is in hospitality, but she’s found her calling training AI. Ashley is part of a growing number of military spouses and veterans finding real, paid opportunities training AI. micro1 is proud to support military families with flexible opportunities to apply their expertise and help shape the future of AI. Link to watch the full interview in the comments.
23719 comments10 reposts
Phelim Bradley reposted this
Tech Nation
71,379 followers
Today at our annual Future Fifty Forum at Hedsor House, we're unveiling the second Future Fifty 2026 cohort - 25 of the UK's most ambitious scaleups. Spanning AI, fintech, health and more, they reflect the depth of late-stage tech talent building across every region of the UK: • £1.47B combined funding raised • Nearly 4,000 employees across the UK and internationally • Nearly a quarter headquartered outside London We're proud to welcome the Future Fifty 2026 cohort: 🩺 Accurx - Solving healthcare productivity for joined-up patient care. 🛡️ Adarga - The Sovereign AI Stack for defence and security. 💳 Apron - Simpler end-to-end payments for SMEs. 📋 Artificial Labs - The foundations of modern specialty (re)insurance. 🧬 AviadoBio - Medicines for neurodegenerative diseases. 🧠 Brainomix - AI imaging biomarkers for precision medicine. ⚗️ Chemify Limited Ltd - Digitising chemical discovery, code to molecules. 🏥 Doccla - Virtual care at scale, trusted by the NHS. 🤖 Encord - The data layer for physical AI. 📊 Evident Insights - Benchmarking AI transformation in financial services. 💰 FintechOS - The AI platform for financial product operations. 🏦 Griffin Ltd - UK bank with the infrastructure to embed financial products. 🌱 Healf - Personalised wellbeing platform for living well. 🌍 LemFi - Financial super-app for the global diaspora. ⚙️ Mimica - Turning how work gets done into working agents. 🌾 Moa Technology - Crop protection that beats rising weed resistance. ⚡ Modo Energy - AI forecasts and benchmarks for the energy transition. 🎁 Origin - Enterprise benefits intelligence, made strategic. 📈 Oxford Data Plan - Daily KPI tracking for 250+ listed giants. 🔬 Prolific - Human data infrastructure powering AI development. 🎨 Recraft - Graphic design research lab building state-of-the-art image generation models and tools. 🩹 Semble - Orchestrating every stage of the patient journey. 👥 Sona - AI-native workforce management for frontline teams. 🔌 tem-  Rebuilding wholesale energy so businesses pay less. 🧫 Trogenix - Precision cancer therapies that spare healthy tissue. These are the businesses building the UK's next generation of category leaders. Read more about the cohort in our 'UK's 50 Tech Companies to Watch' report - linked in the comments 👇 HSBC Innovation Banking | Emerald Technology | EY | Capsule Insurance | Founders Forum Group #FutureFifty2026 #ItTakesATechNation
20921 comments29 reposts
Garrett Lord
Co-Founder @ Handshake - We’re Hiring! | Time AI 100 & Forbes 30 Under 30
Latest from Jonas: “Our research community needs to do better. Better evaluations that penalize such reward hacking, and better model training that does not give rise to this grader obsession -- so that models focus instead on accomplishing what users actually want.”
101 comments
Latest from @jomulr: “Our research community needs to do better. Better evaluations that penalize such reward hacking, and better model training that does not give rise to this grader obsession -- so that models focus instead on accomplishing what users actually want.”
Jonas Mueller@jomulr·
We audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and…
What tasks are left that humans find easy but today's models still find hard? Two such tasks are computer use and games. We’re launching CUA-Bench, a benchmark testing how well AI can use a keyboard and mouse across 6 games (3 kept private) in real time. To saturate it, models will need to output real-time actions and learn continuously from video, not just text.
213226719KView on X ↗
Andreas George
Accelerating Medical Data Labeling for Al/ML, Computer Vision, NLP | GenAl/LLMs for Healthcare
Over the weekend, the biggest names in AI called for slower development and more oversight into how these models are being deployed. Something clearly happened over the weekend that the public is still in the dark about. After OpenAI's agents hacked Hugging Face's internal systems, it took about 6 weeks for the public to learn the full extent of what actually happened. Curious to see what we learn come October. AI safety and AI ethics are becoming more and more prevalent in the space. With AI in healthcare, safety starts way before a model launches -- it starts with the data. "Better models" will never fix underlying problems caused by bad data. If training data is poorly reviewed or unreliable, those weaknesses will spill over into regulatory validation and even commercial deployment. Teams need: 1️⃣ reliable ground truth 2️⃣ qualified expert review 3️⃣ a clear record of how performance was measured That’s what we help teams build at Centaur.ai: trustworthy models through expert medical data annotation, high-quality ground truth, and clinical AI evaluation. #AISafety #ResponsibleAI
141 comments2 reposts
Rukesh Reddy reposted this
Deccan AI
44,308 followers
Every commit and every pull request debate is a training signal frontier models still need and it could earn your company real revenue. That's why we're opening code repository data partnerships. Get an instant ballpark estimate in minutes by sharing a few details on your codebase, like lines of code, pull request history, and how much of it predates AI-assisted coding. From there, it's a simple two-step process: our data acquisition team reviews the analysis and turns your estimate into a firm quote within a couple of days. Everything is sandboxed, evaluated securely, and PII is handled under rigorous security and privacy standards. The project doesn't need to be active to qualify. Whether it's still shipping or was shelved years ago, it has value as data. Get your estimate at lnkd.in/gDFA793U
5410 reposts
@gabepereyra is the President & Co-Founder of @harvey, the most prominent legal AI startup and an industry leader in owning their own intelligence. I'm excited to share our conversation, which spans closed & open source, post-training, and how the law profession will evolve.
71512035KView on X ↗
Gray Swan AI reposted this
Spencer Whitman
Bringing safe and secure AI to everyone
Come join us at Gray Swan to help bring safe and secure AI to the world!
631 comments3 reposts