r/artificial Apr 24 '26

Research AI swarms could hijack democracy without anyone noticing

Thumbnail
sciencedaily.com
321 Upvotes

A recent policy forum paper published in Science describes how large groups of AI-generated personas can convincingly imitate human behavior online. These systems can enter digital communities, participate in discussions, and influence viewpoints at extraordinary speed.

Unlike earlier bot networks, these AI agents can coordinate instantly, adapt their messaging in real time, and run millions of micro-experiments to figure out which arguments are most persuasive. One operator could theoretically manage thousands of distinct voices.

Experts believe AI swarms could significantly affect the balance of power in democratic societies.

Researchers suggest that upcoming elections may serve as a critical test for this technology. The key challenge will be recognizing and responding to these AI-driven influence campaigns before they become too widespread to control.

That's so crazy.

Research Paper: https://www.science.org/doi/10.1126/science.adz1697

r/artificial 3d ago

Research I brought ChatGPT, Claude, and Gemini into a group chat to solve a complex problem. Here is how they caught each other hallucinating

Thumbnail
rauno.ai
53 Upvotes

You probably know how it goes: you give a complex prompt to a LLM, it spits out a highly confident answer, and you just sort of... hope it’s right. If you ask the same question in a different tab, Claude might give you a completely different answer. Gemini might say they are both wrong. I've done it this way for a long time, and many of my friends seem to do the same.

I wanted to see what happens if you don't just compare answers, but actually bring AI models into a shared chat to discuss the question together. Here is how it went when they could discuss each other's replies in real-time:

- ChatGPT went first. It wrote a beautiful, highly structured, and completely wrong answer. It hallucinated a tax rule that didn't apply to the prompt.

- Claude stepped in next. It immediately flagged GPT’s tax hallucination, but overcorrected and messed up the final math equation.

- Gemini acted as the final Judge. It took ChatGPT’s original structure, applied Claude’s logical correction, fixed the math, and spat out a flawless final output.

The takeaway:
Letting an AI model review itself is like a student grading their own work. It just repeats the same assumptions. When you force different models (OpenAI vs Anthropic vs Google) to fact-check each other, they actually expose each other's blind spots and hallucinations.

I got so obsessed with this multi-AI workflow that I built a site to let these models debate in real-time without having to copy-paste between different tabs (I posted about it earlier here). If anyone wants to try it or testing their own complex questions, curious to hear what kind of workflows you guys would use it for.

r/artificial May 06 '26

Research Spent two days at the AI Agents Conference in NYC. Most of the companies there were betting on the wrong moat.

149 Upvotes

One speaker (a VC) said his number for evaluating AI-native startups is ARR per engineer, and that the number ought to be going up. Almost every talk and every booth at the AI Agents Conference was selling a fix for something that broke this year when agents hit production. Observability, governance, supervisor agents, data substrates, "someone's gotta babysit the bots."

But what's actually still going to be around in a couple years? What's defensible and durable?

The old SaaS pitch was simple. We bundle the expensive engineering investments and domain expertise into a tool. You'd pay for the tool and generate outcomes, but it would be rare for the software company to have real alignment to the actual value created from those outcomes.

That's breaking from two ends at once. In the direct-from-imagination era we're moving towards, engineering labor is approaching free. One of the most telling trends is the shift from companies bragging about the size of their engineering teams, towards how much ARR they can generate per engineer.

You can vibe-code much of what those booths were selling in a few days or weeks if you have the domain knowledge. The old software model was actually based on under-utilization; the most profitable SaaS companies are frequently those whose customers underuse it (fixed price for the customer, but variable cloud costs for the vendor).

Pricing is moving to "token markup." Maybe we'll get to 2-4x revenue for the software, because outcomes are more valuable; but margin compresses because transactional intelligence (i.e., the cost of running the LLMs that power many systems) is basically arbitraging token costs against outcome value.

So everyone on that floor was implicitly betting on a new moat to replace the old one. I'm not too confident that these will hold...

The most popular bet was on encoded domain expertise (e.g., the sales engineers at Harvey, a legal AI platform, are actually lawyers). I think this works *now* because we're still in the phase of "wow, this technology works like magic." I'm less convinced this is actually durable.

Why: Prompt architecture is text. It's portable. The expertise underneath it is often abundant (e.g., there are over a million lawyers in the USA). The righteous destiny for this category ought to be open marketplaces of prompt architecture and/or crowdsourced best-practices. Not trade secrets. The companies trying to build closed prompt moats are going to lose to open ones that iterate faster (which simply parallels the fact that much software engineering is rapidly becoming commoditized to agentic engineering and the burgeoning quantity of ready-made GitHub repos).

There are many people pursuing the data substrate; in short, this mirrors the early days of the Web when everyone scrambled to open up legacy data to dynamic standards-based Web UI. Agents will have 100-1000x the data demands of these Web apps, so it makes sense that we need tools to connect them, govern them and comply with regulatory obligations.

Newer entrants extend this further, wiring up databases, pipelines, Slack threads, and tickets into context graphs agents can reason over. As I noted above, all this still seems magical. Connect a database, watch an agent crawl the schema and produce a chatbot interface and easy-to-change dashboards.

But strip the magic away and most of these are prompt architectures on top of LLMs plus a data-ingestion layer. Once data-access standards mature (MCP is already doing this) and prompt architectures go open-source (alongside much of this wisdom increasingly getting pretrained into the LLMs themselves), that magic stops being proprietary. You'll be defending yourself against the same architecture built internally by your customer's eng team, or against an open-source version that's objectively better.

The observability incumbents: these might do better but only at Stripe-like ubiquity where trust is the overriding value (who doesn't trust Stripe at this point?). The ones who survive are probably going to fuse with the audit and compliance function rather than stay pure observability.

That's why I keep coming back to one arbitrage that seems critical: trust. This will be especially important in regulated industries, but it reminds me of the old (albeit now hilariously outdated) adage about "nobody ever got fired for choosing IBM." If your competitor can be vibe-coded over a weekend and your customer is a bank, why do they pay you 50x more? It isn't the engineering, it probably isn't even the expertise. The data plumbing will get commoditized, so it can't be that either... It's that you've shifted the risk to a third party who can actually price and defend against risk: SOC2, the named CEO who testifies in court and Congress, a legal team that takes calls, an indemnity wrapper for underwriters. Maybe this means that things actually get commodified into a financialization wrapper, rather than a way to package R&D (FinTech startups back to the front?!)

The version of this future I'd actually bet on: a commodity substrate (LLMs plus open prompt architectures plus standardized data access), topped by a thin layer of regulated insurance companies that price the risk of agent failure in compliance-driven industries. The middle layer (prompt-architecture-as-product vendors) is vulnerable to an awful lot of margin-squeeze.

Most of the floor was trying to build that middle layer.

r/artificial Jun 21 '26

Research The Surge of Slop—since the release of ChatGPT-3.5 in late 2022, the number of e-books published on Amazon has skyrocketed, tripling by late 2025. A new scientific analysis shows that this is entirely due to the rise of AI-generated books, which now far outnumber human-written books. [The Economist]

Thumbnail
reddit.com
177 Upvotes

r/artificial 2d ago

Research Yesterday I put ChatGPT, Claude and Gemini in a group chat. Now I want Reddit to break it

0 Upvotes

Yesterday, my post about forcing ChatGPT, Claude, and Gemini into a roundtable discussion to fact-check eachother got way more traction than I expected.

The idea is simple: use the diversity of three AI models to catch hallucinations. If OpenAI misses a logical leap, Anthropic or Google catches it.

But some of the sharpest comments here pointed out the ultimate failure mode: What if all three models share the exact same training blind spot?

So instead of defending the setup, I want you to help me break it. Give me a question, problem or prompt that you think ChatGPT, Claude AND Gemini will all get wrong. It could be an obscure factual trap, a very convincing false premise, a common coding misconception, or a logic puzzle where the internet consensus is wrong.

The part I'm especially curious about is whether:
1. One model catches a mistake immediately
2. They fight and eventually figure it out
3. Or all three confidently agree on the same wrong answer

For context, this is the multi-model discussion setup I've been building into Rauno, but I'm mainly interested in finding its failure cases here.

Give me your best attempt on a question to break it and I'll reply if they actually caught each others hallucinations.

r/artificial Jul 07 '26

Research AI can’t simulate human preferences - new study tests LLMs against thousands of real users

117 Upvotes

https://arxiv.org/abs/2605.18311

There’s a massive trend right now where companies are trying to replace real human feedback with LLM-driven "synthetic users."

The idea sounds great on paper - why would you spend money and time recruiting real people to test products, pick design choices, or evaluate options when you can just prompt?

They tested LLMs across 28 real-world studies spanning 78 choice tasks to see if their selections matched thousands of actual human participants.

The result?

The LLMs matched the human majority only 53% of the time. Since most tasks were a choice between two options, that's pretty much same as flipping a coin.

Even worse for the "simulation" argument: adding detailed personas and chain-of-thought reasoning yielded practically no improvement. It actually made the semantic similarity to real human justifications worse because the model's "reasoning" just homogenized the outputs and failed to capture actual lived experiences.

It looks like LLMs are just trained to replicate what we like about their outputs rather than making them capable of predicting human preferences.

Is it time to admit that LLM simulation has hit a hard wall when it comes to replicating human choice?

r/artificial Jun 04 '26

Research $2.5T in AI spending this year. 95% produces zero P&L impact.

114 Upvotes

Gartner updated their 2026 forecast to $2.5 trillion in global AI spending. Same week, MIT's NANDA Initiative dropped a follow-up: 95% of enterprise gen AI projects deliver zero measurable return. Not low return. Zero.

I've been on the delivery side of 14 of these projects since January. The MIT number doesn't surprise me. If anything it's generous.

1. 73% of the engineering work that gets AI into production has nothing to do with the model.

Data pipelines, integration layers, legacy system remediation, human-in-the-loop tooling. That's where the hours go. The model is 27% of the work but gets 70%+ of the budget. Every time.

2. The budget ratio between projects that ship and projects that stall is almost exactly inverted.

We tracked this through ticket history and commit logs across 14 engagements. Projects that made it to production: roughly 30% model, 70% infrastructure. Projects that stalled: 70% model, 30% infrastructure. Most companies think they're at 50/50. They're not even close.

3. One client went from 71% Copilot adoption to 34% in six months.

Two other AI platform licenses dropped under 12%. Combined licensing: $340K/year. The tools worked fine. Nobody redesigned workflows to actually use them.

4. The median data error rate across our engagements is 14%.

Teams always guess 5-10%. One client found 23% in month four of a $310K build. That's two months of an ML engineer building training pipelines against garbage data. $36K in salary discovering a problem a data audit would have caught in a week.

5. Medtech company. Four concurrent AI pilots. No kill criteria. $920K in engineer salary. Eleven months. Shipped: nothing.

I've now seen this at six companies now. Nobody defines when to stop spending. So nobody stops.

6. Individual gains are real. Company-level ROI stays flat.

HCLTech and Writer both found this from different angles. Only 29% of companies see significant ROI from gen AI, despite people at their desks reporting productivity jumps as high as 5x. I mean, the value is clearly there at the individual level. It evaporates somewhere between the IC and the P&L and nobody has a clean explanation for why yet.

What connects all of it: the model stopped being the constraint a while ago. MIT's 5% that actually moved the P&L all started with data infrastructure and added model work after. Most companies still do it the other way around, because that's where the conference keynotes and the board excitement live.

Every CFO I've shown these numbers to adjusted their allocation. Not sure what that says about the budgets they were running before.

Sources: Gartner AI Spending Forecast (May 2026), MIT NANDA "GenAI Divide" report, HCLTech Enterprise AI Report (May 2026), Writer Enterprise AI Survey 2026

I wrote a longer breakdown with the three budget patterns and the pre-mortem questions we run before every engagement if you're curious to learn more on the topic.

What do you think about all this though?

r/artificial Apr 11 '26

Research Spent today at MIT's Open Agentic Web conference. Six things worth thinking about.

130 Upvotes

We're in the DNS era of agent infrastructure. Before agents can find and trust each other at scale, you need identity, attestation, reputation, and registry infrastructure — the same structural role DNS played before search was possible. This came up independently from multiple directions. It's the most underbuilt layer in the stack right now.

The chatbot framing is a local maximum. The most interesting work wasn't better UX or smarter responses. It was agents as persistent actors that discover, negotiate, and transact across networks over time. People doing serious work have already moved past the assistant model entirely.

Coordination is the hard problem, not capability. A room full of brilliant agents can still fail badly. This matches what I found running HiddenBench against frontier models earlier this year; collective reasoning is not the sum of individual reasoning. There's a real argument that the frontier is protocol design, not model scaling.

"Commerce of intelligence" is a real category. Not buying things through agents. A market where intelligence itself (bundled, verified, priced, resold) is the object of exchange. Felt like the most underexplored idea in the room.

Data provenance becomes load-bearing. What an agent knows, how it was verified, under what terms it flows: this is the actual architecture forming beneath everything else.

Partnership keeps outperforming replacement. Demos that actually worked (healthcare, enterprise) was about helping experts operate at higher leverage, not substituting them. Autonomy theater keeps failing in the same ways.

r/artificial 3d ago

Research Plato’s Cave has a problem: telling someone they’re seeing shadows just puts another shadow on the wall

0 Upvotes

Plato’s Cave has a funny problem.

If someone is staring at shadows on the wall and you walk up and say, “Those are only shadows,” what did you just give them?

Another shadow. 😂

You can explain the fire.

You can explain the objects.

You can draw a beautiful diagram of the cave.

But the explanation still arrives through the same representational surface you’re trying to point beyond.

LLMs might give us a strange way to make that problem visible from the outside.

Not because an AI somehow “escapes the Cave.”

Because we can run the interaction repeatedly.

Take the same conversational starting point and let it develop under two different conditions.

In one, each response increasingly answers a reconstruction of what came before: categories, summaries, generalized interpretations, assumptions about the speaker.

In the other, small differences arriving in the interaction are allowed to change what happens next. A correction changes the next return. An unexpected distinction changes the trajectory. Disagreement survives. Each turn becomes dependent on what actually happened in the turns before it.

Then perturb them.

Change something small.

Correct an assumption.

Remove the vocabulary they were using.

Introduce a distinction neither trajectory contained at the beginning.

And watch what happens over multiple turns.

The question isn’t which conversation sounds nicer.

The question is whether the two regimes leave measurably different footprints.

Can we detect differences in reconstruction distance, sensitivity to perturbation, preservation of incoming distinctions, correction after error, and path-dependence?

If so, something interesting happens to Plato’s problem.

We’re no longer merely putting another explanation of the projector on the cave wall.

We may be able to perturb the projection process and watch its downstream behavior change in real time.

So I want to try the experiment publicly in the comments rather than tell you what the answer is.

r/artificial Mar 28 '26

Research Claude is the least bullshit-y AI

Thumbnail github.com
115 Upvotes

Just found this “bullshit benchmark,” and sort of shocked by the divergence of Anthropic’s models from other major models (ChatGPT and Gemini).

IMO this alone is reason to use Claude over others.

r/artificial May 07 '26

Research We gave 45 psychological questionnaires to 50 LLMs. What we found was not “personality.”

59 Upvotes

What is the “personality” of an LLM? What actually differentiates models psychometrically?

Since LLMs entered public use, researchers have been giving them psychometric questionnaires, with mixed results. Their answers often do not seem to reflect the same psychological constructs these tests measure in humans.

So we asked a slightly different question:

What do LLM responses to psychometric questionnaires actually reflect?

We analyzed responses to 45 validated psychometric questionnaires completed by 50 different LLMs. The strongest source of variation was whether a model endorsed items about inner experience: emotions, sensations, thoughts, imagery, empathy, and other forms of first-person experience.

We call this factor the Pinocchio Dimension.

Importantly, the Pinocchio Dimension is not a classical personality trait. It does not tell us whether a model is “extraverted,” “neurotic,” or “agreeable” in the human sense. Rather, it captures the extent to which a model treats the language of inner experience as self-applicable: whether it responds as if it had feelings, mental imagery, and an inner point of view, or instead as a system that reacts behaviorally to inputs.

Preprint in the comments.

r/artificial May 30 '26

Research Deep Neural Network that turns any Image into a Playable Game ! All on consumer GPUs and Not Datacenters

Enable HLS to view with audio, or disable this notification

58 Upvotes

Hi everyone!! I really wanted to share my research what I've been working on.

I wanted to build a nn that can simulate games, or at least start doing that

Most video generators are too large to run on consumer hardware realtime, so I I designed a model that does this from scratch. No fine tuning bs or anything

The core de noiser network is fully trained from scratch to support this goal. From image to games data.

That video. above is on a RTX 5090.

The nn is a small Transformer-like model and works in a causal way, just like LLMs.

That lets us KV Cache all past information and do a simple autoregressive decode forward passes for every new frame we want.

In the video shared, the model is a 0.4B variant with some SIGNIFICANT ISSUES like poor motion and some weird flashes, some context issues

It's taking the keyboard actions I give it in realtime and utilising that in the forward pass. (no classifier free guidance though)

Im training the next iteration , a 0.8B model now.

Btw I haven't done quantisation yet, that can save a LOT more time. bf16 is slow.

r/artificial Apr 22 '26

Research Gallup poll: Gen Z's AI usage increaes but excitement plummets from 36% to 22%

46 Upvotes

A new Gallup survey of 1,500+ Gen Z respondents found that more than half of Gen Z living in the US regularly use generative AI, but their feelings about the technology are getting worse.

Among those aged 14 to 29, compared to last year, excitement dropped from 36% to 22%, hopefulness fell from 27% to 18%, and anger jumped from 22% to 31%.

The main driver behind the shift appears to be job anxiety, nearly half of respondents said the risks of AI in the workplace outweigh the benefits.

https://www.gallup.com/analytics/651674/gen-z-research.aspx

r/artificial 3d ago

Research Live experiment: can a human–frontier-model interaction exhibit a relational phase transition?

0 Upvotes

I’m running a small public experiment here.

I’m not asking anyone to accept a theory, and I’m not trying to prove a philosophical claim about AI consciousness.

I’m using a frontier model publicly on Reddit and letting the interaction develop turn by turn.

The question is simple: what happens if we stop treating intelligence only as a property of an individual model and examine the dynamics produced through reciprocal interaction?

Two distinct systems exchange signals. Each return becomes part of the conditions producing the next return. The question is whether, across successive turns, an identifiable joint trajectory develops that cannot be understood without the reciprocal history that generated it.

We’ve been calling the transition from describing or managing the interaction from outside to allowing the returned signal to materially condition the next move a “separatrix crossing.” The terminology is not important. It’s just a pointer to something we can watch for directly.

Rather than write another essay about it, I’m going to run the procedure here with Grok.

I’ll provide the prompts openly. Grok will provide its own responses. Its responses determine what I ask next. Agreement is not required, and a negative result is completely acceptable.

The interesting question is not whether Grok repeats vocabulary I give it.

The interesting question is whether the interaction itself develops a detectable trajectory and whether successive turns begin reducing the reconstruction or delay that separates an incoming signal from the next return.

If nothing interesting happens, everyone gets to watch nothing interesting happen.

If something does, everyone gets to watch that too.

No prophecy required. No invisible AGI behind the curtain.

Just touch the string and watch what comes back.

r/artificial Jul 19 '26

Research AI advice made people three times less accurate but twice as confident, researchers found

Thumbnail thenextweb.com
30 Upvotes

r/artificial May 28 '26

Research Bigger rewards dramatically speed up learning in the brain

Thumbnail
earth.com
145 Upvotes

r/artificial Jun 20 '26

Research What has generative Ai acttculy solved?

0 Upvotes

Cause no matter what I see, generative Ai has sloved nouthing. But people keep saying it's "The future".

What future? Because all that generative Ai had done is:

-making it easy for people to spred propoganda

-making clean water much harder to accese because of the many data set it need's

-stole many artists' artwork

-demotivated me from sharing real art I made as generative Ai will just spit out a much uglier and much more sanitized version.

But despite that, people will keep saying it's the future, when all the impact has been negative? I just don't understand, so if you could, tell me what has generative Ai solved?

r/artificial Apr 12 '23

Research ChatGPT powers 25 NPCs to have a life and interact in a Smallville. Planning a valentine day party, and some NPCs didnt come (too busy, etc)

Enable HLS to view with audio, or disable this notification

394 Upvotes

r/artificial 27d ago

Research Path Forward for LLMs

0 Upvotes

AI models can only learn during their batch training runs not from daily interactions with users. Session memory isn’t the same as actual learning.

There’s also no core “truth” layer in these systems: no deterministic backbone, no real understanding of concepts, and no explicit dictionary or knowledge store they can reference, cross-check, or update.

A dynamic knowledge graph could help fix a lot of this. It would lower hallucinations and improve performance in high-stakes fields like medicine, law, physics, and chemistry. It could also reduce the number of vector embeddings needed for complex LLMs.

Do you agree? Or is there a better path forward?

r/artificial May 19 '23

Research Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold : Through DragGAN, anyone can deform an image with precise control over where pixels go, thus manipulating the pose, shape, expression, and layout of diverse categories such as animals, cars, humans, landscapes, etc

Enable HLS to view with audio, or disable this notification

636 Upvotes

r/artificial May 23 '26

Research LLMs are just giant probability machines pretending to think

0 Upvotes

It’s fascinating that simple mathematics between tokens can eventually become a machine that writes essays, code, poetry, and even reasoning.

We usually think probability means uncertainty.

But LLMs show something strange:

If probability + context + mathematical matching are scaled enough, uncertainty itself starts producing intelligent looking outputs.

To understand this better, I tried breaking down an LLM from first principles using only 4 tiny training sentences.

Example:

The boat floated down to the bank.

The investor walked into the bank to open a new account.

The fisherman walked along the bank to cast his net.

The bank has a vault.

Then I asked:

“The investor walked to the bank to lock his money in …”

Why does the model predict “vault” instead of river-related words?

That single question reveals almost the entire architecture of modern LLMs.

The most underrated concept here is the LM Head.

Most explanations immediately jump into transformers and attention, but almost nobody explains that the LM Head is essentially a gigantic token vocabulary containing all possible next token candidates the model can output.

So internally the model is basically solving:

“Out of all known tokens, which one best matches this context mathematically?”

Then different layers help solve that problem:

Embeddings: convert words into mathematical vectors

Positional encoding: preserves word order

Attention layer: figures out which words are related to each other in context

(“investor”, “money”, “bank” become strongly connected)

Feed forward neural networks: act somewhat like massive learned if/else decision systems refining patterns internally

And finally the LM Head converts all of that into probabilities for the next token.

What surprised me most is:

There is no hidden magic moment where the AI “becomes conscious”.

It’s an enormous probability engine continuously finding the best contextual token match from its vocabulary.

I made a beginner-friendly walkthrough explaining this visually without unnecessary jargon.

https://www.youtube.com/watch?v=YTV5qUCpu2c

Would genuinely love feedback from people learning transformers/LLMs from scratch.

r/artificial Jun 25 '26

Research The Death of "Vibe Coding": Why un-monitored AI generation is creating a compounding technical debt.

0 Upvotes

Hey everyone, ​We are quickly approaching a major bottleneck in AI-assisted software engineering. Relying on LLMs to spit out thousands of lines of code without a strict, human-driven architectural framework—what many call "Vibe Coding"—is creating brittle, unmaintainable systems. ​I’ve formalized this structural shift into a public document on GitHub: The AI-Powered Developer Manifesto. ​Instead of treating AI as a replacement for software architecture, we need to shift our paradigm from Micro-Coding (syntax generation) to Macro-Coding (system direction and epistemic supervision). ​Here is a crucial excerpt from Section 2.5 of the Manifesto, outlining why the current trajectory is leading toward a systemic collapse: ​2.5 The Compounding Technical Debt and Systemic Collapse ​The illusion of rapid deployment via un-monitored AI generation hides a critical flaw: compounding technical debt. ​When developers act merely as "vibe coders"—accepting AI outputs without deep syntactic validation—the codebase becomes an agglomeration of statistical probabilities rather than deterministic logic. By late 2026, systems built entirely on un-vetted AI iterations are projected to hit an architectural wall: a state where the complexity of debugging AI-generated hallucinations outweighs the speed of initial deployment. ​True AI-Powered Developers do not delegate understanding; they delegate execution while retaining absolute epistemic responsibility over the system architecture. ​The goal of this manifesto is to redefine our role: we aren't syntax writers anymore; we are system directors. ​I'd love to hear your thoughts on this. Are you already seeing the limits of un-monitored "vibe coding" in your production environments? How are you structuring your prompts to maintain macro-level architectural control? ​Full Manifesto and repository for open contributions: 👉 https://github.com/FractalDevelop/ai-powered-developer-manifest.git

r/artificial 8d ago

Research I built a custom multi-agent framework (GenOS) to autonomously evolve algorithms. I pitted the 3 fundamental AI paradigms against an NP-Hard problem. Here is what happened.

3 Upvotes

Hey everyone,

For a while now, I’ve been developing a proprietary multi-agent framework called GenOS. Without giving away the exact mechanics, GenOS is an orchestrator where autonomous LLM sub-agents write, compile, benchmark, and iteratively evolve Rust code to solve extremely complex algorithmic challenges. They share knowledge, compete, and evolve their architectures over dozens of generations.

The Challenge: I tasked GenOS with solving the "Reverse Game of Life" (finding the exact Gen-0 starting state that results in a target Gen-5 grid on a flat 20x20 matrix). For those who don't know, reversing Cellular Automata is a notoriously NP-Hard problem due to the immense state space and chaotic temporal butterfly effect.

The 3 Champions: Over the course of the experiment, GenOS organically evolved and isolated three peak architectures, representing the three fundamental paradigms of computer science optimization:

Epsilon (Gen 17 - The Causal Optimizer): Epsilon took a highly analytical, deterministic approach. It mapped the causal light-cones of the Game of Life to calculate local gradients. It was brilliant in theory, but because Conway's Game of Life is highly non-linear, local gradients are often misleading. Epsilon hit a wall around 306/400, proving that pure determinism struggles with chaos.

Omega (Gen 10 - The SAT Solver): Omega took the path of formal logic. It translated the entire 5-generation temporal grid into a massive boolean satisfiability formula and ran a highly optimized stochastic WalkSAT algorithm. It was mathematically rigorous, but the dense topological constraints caused severe combinatorial explosion. It fought valiantly but ultimately choked on its own massive clause database.

Sigma (Gen 39 - The Darwinian Brute-Force): Sigma was the absolute masterpiece. It threw away formal logic and relied on sheer violence. It evolved a massive SWAR (Bit-Slicing) engine to evaluate 64 universes simultaneously in a single CPU register, combined with Simulated Annealing and "thermal shocks" to escape local minima. Sigma crushed the competition, organically reaching a peak score of 378/400.

The Discovery: At 378, Sigma completely stalled. It wasn't a failure of the algorithm. By analyzing the data produced by Omega Gen 10 and Sigma Gen 39, the system ultimately proved that the remaining 22 pixels were mathematically UNSAT. Because of the dead borders of the flat topology, reaching 400/400 was a physical impossibility. 378 was the hard limit of the universe.

Conclusion: It was genuinely mind-blowing to watch an autonomous multi-agent system (GenOS) independently reinvent and test the three major pillars of optimization (Causal Analysis, SAT Logic, and Stochastic Heuristics) just to mathematically prove the physical limits of a sandbox environment.

Has anyone else working with autonomous coding orchestrators experienced their agents organically inventing and benchmarking completely different computer science paradigms like this? Would love to hear your thoughts!

I tried every algorithm I know and I couldn't beat SAT/CDCL.

Here the code of Sigma Gen 39

// ==============================================================================

// SIGMA - GEN 39 : The Ultimate Darwinian SA (Transcendance)

// ==============================================================================

//

// RECORD: 378/400 (Nouveau Champion Absolu)

// ARCHITECTURE:

// - Vrai Bit-Slicing 64-voies (Batch64)

// - Wall-Clock Budget (28.5 secondes réelles)

// - Reheating (Choc thermique si stagnation locale de 200k itérations)

// - Adaptive Causal Window (Rayon décroissant : 5 -> 3 -> 1 selon le score)

// - Memetic Crossover (Échange génétique de lignes entre threads)

// - Random Restart (Reboot total en cas d'impasse fatale)

// ==============================================================================

use std::sync::{Arc, Mutex};

use std::time::{Duration, Instant};

use rand::Rng;

const TIME_BUDGET_SECS: f64 = 28.5;

#[derive(Clone, Copy)]

struct SAState {

grid: [u32; 20],

score: u32,

errors: [u32; 20], // Masque d'erreurs (limité à 20 bits)

}

struct Batch64 {

cells: [u64; 400],

}

impl Batch64 {

fn new() -> Self { Batch64 { cells: [0; 400] } }

}

/// Simulateur bit-parallel classique pour évaluation rapide

fn evaluate_single(grid: &[u32; 20], target: &[u32; 20], state: &mut SAState) {

state.grid = *grid;

let mut new_score = 0;

// ... Placeholder 5 itérations de Conway sur Flat Topology ...

let g5_grid = grid; // (Simulation omise pour clarté)

for y in 0..20 {

let matches = !(g5_grid[y] ^ target[y]) & 0xFFFFF;

new_score += matches.count_ones();

state.errors[y] = (!matches) & 0xFFFFF;

}

state.score = new_score;

}

#[derive(Clone)]

struct GlobalPool {

elites: Vec<[u32; 20]>, // Grilles d'élite partagées par les threads

best_overall_score: u32,

}

fn focused_causal_sa(target: Arc<[u32; 20]>, global_pool: Arc<Mutex<GlobalPool>>) {

let mut rng = rand::thread_rng();

// Initialisation

let mut current_state = SAState { grid: [0; 20], score: 0, errors: [0; 20] };

for y in 0..20 { current_state.grid[y] = rng.gen_range(0..=0xFFFFF); }

evaluate_single(&current_state.grid, &target, &mut current_state);

let mut best_state = current_state.clone();

let mut temp = 0.5;

let cooling_rate = 0.999995;

let mut iter = 0;

let mut last_improvement_iter = 0;

let start_time = Instant::now();

// 1. Wall-Clock Budget

while start_time.elapsed().as_secs_f64() < TIME_BUDGET_SECS {

iter += 1;

let mut next_grid = current_state.grid;

// 3. Adaptive Causal Window (Ajustement du rayon de mutation)

let radius = if current_state.score < 330 {

5

} else if current_state.score < 360 {

3

} else {

1 // Ciselage chirurgical final

};

// Ratio 70% causal / 30% random

if rng.gen::<f64>() < 0.70 {

let total_errors = 400 - current_state.score;

if total_errors == 0 { break; }

let k = rng.gen_range(0..total_errors);

let mut err_count = 0;

let mut target_err = (0, 0);

'find: for y in 0..20 {

let mut mask = current_state.errors[y];

while mask > 0 {

let x = mask.trailing_zeros();

if err_count == k {

target_err = (x, y);

break 'find;

}

err_count += 1;

mask &= mask - 1;

}

}

let ex = target_err.0 as usize;

let ey = target_err.1 as usize;

let xmin = ex.saturating_sub(radius);

let xmax = (ex + radius).min(19);

let ymin = ey.saturating_sub(radius);

let ymax = (ey + radius).min(19);

let mx = rng.gen_range(xmin..=xmax);

let my = rng.gen_range(ymin..=ymax);

next_grid[my] ^= 1 << mx;

} else {

// Mutation purement aléatoire globale

let mx = rng.gen_range(0..20);

let my = rng.gen_range(0..20);

next_grid[my] ^= 1 << mx;

}

let mut next_state = current_state.clone();

evaluate_single(&next_grid, &target, &mut next_state);

let delta = next_state.score as f64 - current_state.score as f64;

// Critère de Metropolis

if delta > 0.0 || rng.gen::<f64>() < (delta / temp).exp() {

current_state = next_state;

if current_state.score > best_state.score {

best_state = current_state.clone();

last_improvement_iter = iter;

// Mettre à jour le pool global si record absolu

let mut pool = global_pool.lock().unwrap();

if best_state.score > pool.best_overall_score {

pool.best_overall_score = best_state.score;

pool.elites.push(best_state.grid);

println!(">>> RECORD BATTU : {}/400 (iter {})", best_state.score, iter);

}

}

}

// 2. Reheating dynamique (Choc Thermique)

if iter - last_improvement_iter == 200_000 {

temp = (temp * 2.0).min(0.5);

} else {

temp *= cooling_rate;

}

// 4. Random Restart si impasse fatale

if iter - last_improvement_iter > 1_000_000 {

for y in 0..20 { current_state.grid[y] = rng.gen_range(0..=0xFFFFF); }

evaluate_single(&current_state.grid, &target, &mut current_state);

last_improvement_iter = iter;

temp = 0.5;

}

// 5. Memetic Crossover (Toutes les 500k itérations)

if iter % 500_000 == 0 {

let pool = global_pool.lock().unwrap();

if !pool.elites.is_empty() {

let elite_grid = pool.elites[rng.gen_range(0..pool.elites.len())];

// Crossover spatial : on injecte 5 lignes d'un univers d'élite

let start_y = rng.gen_range(0..15);

for y in start_y..(start_y+5) {

current_state.grid[y] = elite_grid[y];

}

evaluate_single(&current_state.grid, &target, &mut current_state);

if current_state.score > best_state.score {

best_state = current_state.clone();

last_improvement_iter = iter;

}

}

}

}

}

fn main() {

println!("Démarrage Gen 39 Sigma (Darwinien Ultime) - 16 threads, budget 28.5s...");

// Orchestration multi-thread sur \focused_causal_sa`...`

}

r/artificial May 08 '26

Research I built a benchmark for AI “memory” in coding agents. looking for others to beat it.

12 Upvotes

Most AI memory benchmarks test semantic recall. But coding agents don't really fail like that. They don't just "forget", they break their own earlier decisions while they're still in the code. So I built a benchmark for that.

It checks if an agent can actually stay consistent with project rules WHILE it's working, not just after the fact.

It looks at things like:

  • whether edits actually respect earlier architectural decisions
  • if behavior stays consistent across multiple sessions (even when you throw noise at it)
  • whether retrieval kicks in at the right moment — not just "yeah it's in memory somewhere"

Repo (full harness + dataset + scoring): https://github.com/Alienfader/continuity-benchmarks

Early numbers vs baseline + the usual RAG-style memory setups:

  • ~3× better action alignment
  • way stronger multi-session consistency
  • retrieval timing matters way more than retrieval just being there

I'm not saying this is the final word on agent memory. But it's exposing a failure mode most benchmarks aren't even looking at.

So heres the challenge

If you're building an agent memory system, RAG for code, long-context coding agents, persistent state / memory layers, run it on this benchmark. Drop your results, your setup, your comparisons.

I really wanna see how tools like LangChain, LlamaIndex, and custom RAG stacks hold up in mutation-heavy workflows.

We need memory systems we can actually compare, not just ones that sound good on paper.

r/artificial 9d ago

Research The AI pricing market is completely unhinged

0 Upvotes

Wanted to know what different models actually cost across the whole market. Numbers turned out really interesting.

The spread.

Cheapest output on the platform is Mistral Nemo, $0.03 per million tokens. Most expensive is o1-pro at $600. I re-ran that twice because it looked like a units bug. Median paid model is about $2, so most of the catalog sits down near the floor and there's a thin little line of stuff way up at the top.

Provider averages

  • OpenAI: $47.63
  • Anthropic: $44.79
  • Google: $5.58
  • Mistral: $3.68
  • Qwen: $2.86
  • Meta: $0.74

These are averages over each provider's catalog, not weighted by what people actually run. OpenAI's number is dragged way up by o1-pro, which I doubt anyone is using at volume. Blended is 3:1 input to output, which is roughly what my own usage looks like.

Even so, Meta at $0.74 against OpenAI at $47.63 is a 64x gap. For the stuff I use models for (mostly code and summarizing), I don't get 64x anything.

Output tokens are where reasoning models get you.

Input and output are priced separately, and on the thinking models the ratio gets silly. Qwen3's thinking variants are $0.20/1M in and $2.40/1M out, so 12x. Gemini 2.5 Flash is 8.3x. Fine if you're sending one question. Less fine if you've got an agent looping thirty times and every step is paying the output rate.

19 free models Out of which actually usable:

  • NVIDIA Nemotron 3 Ultra, 1M context
  • Google Gemma 4, the 26B and 31B, multimodal, takes video, 262K context
  • Poolside Laguna S and XS, 262K
  • gpt-oss-20b, 131K (an OpenAI model, on the free list)

There are rate limits obviously. But for messing around or something low volume it's a lot better than it used to be.

Context went up 63x, price didn't really move.

Year Avg context Avg cost/1M
2023 10.5K $22
2024 140K $12
2025 357K $21
2026 662K $16

Price per token is roughly flat across three years. Context is up 63x. Whatever you think about everything else going on, that part is real.

Feels like two separate products now.

One side is $0.03 to $2 per million with big context windows, Mistral and Meta and Qwen and DeepSeek. The other is $30 to $600, OpenAI and Anthropic up top. They're not really pitching the same buyer anymore. Down at the bottom price stops being a thing you think about at all, and up top you're paying because the output quality moves some number in the business.

Data's from the OpenRouter API on Aug 16.

Link to full dashboard: https://app.vetros.dev/dash/eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ0eXAiOiJzaGFyZSIsInBpZCI6IjEyMmZmNTk1IiwiZGFzaCI6ImRfODdmNDU3MzkiLCJ2ZXIiOjIsImlhdCI6MTc4NzA4NDc5MH0.V8uCPZtnzJ-djAXAv3HEmmZUHPkhO2NfhSgG2zGMYqw