Rendered at 09:05:56 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
tristanj 13 hours ago [-]
Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data.
I enabled all data sharing settings but still don’t have a message about free use on that screen - the help page says free tokens are available to “some” users - is that 1% of users, 40% of users, etc?
Does your screen have the message that you’re getting free tokens?
I tried following this page, and it's certainly a lot more complex than what Meta is offering. Different price tiers, opt-in configurations, usage based availability.. I'll take the 10x discount for flipping a param switch over this all day long.
stingraycharles 7 hours ago [-]
Seems pretty clear to me: enable it for the projects you want, and there’s a 1M / 10M token limit per day, depending on the model you use. Assuming an average context size of 100k tokens, that is 10 to 100 requests, which is not a lot. Reason enough to prefer actually paying for Meta as well.
no-name-here 6 hours ago [-]
Does it show the free tokens message for you after you enable sharing? It does not for me.
Also, in a desktop browser at the page's [1] lower left it says "Personal organization", so if the 'organization' term is the concern, OpenAI still seems to use the term 'organization' even for personal accounts.
I actually really like that pricing strategy. It's very transparent
bdangubic 9 hours ago [-]
I love the idea of this pricing strategy but there is no way meta is not training on your data regardless of your monthly invoice
simonw 8 hours ago [-]
So you think the only difference between the $1.25/million token plan and the $0.10/million token plan is that you pay them more to both lie to you and breach their contractual obligation to you?
bdangubic 6 hours ago [-]
when has Meta ever not broken their contractual obligations (I am being serious here)? are we seriously discussing/expecting any sort of privacy related to Meta?
you can pay whatever they want, they will train and use your data, I figured this is not something that should be discussed but obviously I have been mistaken...
JimDabell 5 hours ago [-]
> when has Meta ever not broken their contractual obligations (I am being serious here)?
If you are being serious, then you have a wildly distorted view of the world. No organisation can routinely break all of their contractual obligations. If you think Meta are doing this then you are not seeing Meta, you are seeing a fictional bogeyman.
simonw 6 hours ago [-]
This is one of those situations where we are both completely baffled by the position held by the other person.
The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.
ben_w 6 minutes ago [-]
> The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.
To me, their history suggests that they know more than most about how to get away with breaking both the spirit and the letter of the rules.
I would not know which to expect in any given instance.
However, even if they were to surprise me by being perfectly meticulous about the letter of the rules from now on, I have so little trust in them that I would expect some technicality somewhere in the language of the contract.
teiferer 3 hours ago [-]
How does money change that trust? They certainly have breached their word on this in the past (for non-paying users of Facebook).
discordance 28 minutes ago [-]
Call me childish but it was worth a shot...
"Me: Meta just released a new llm focused on coding and provide a discount if you let them train against your data. I don't like Meta and I think they are a net negative in our world. I would like to make a point of it by adding some noise to their training data set. Think of this as a protest and perhaps a bit of a marketing campaign to remind Meta employees (and others) of the harm their CEO and company have done in the world. To the problem... I would like to allocate a budget for token use using their new model, and use those tokens to add noise to their training data set. This is a coding model and my initial thoughts are to ask it to solve typical CS and common programming related problems but then give Muse feedback that guide it towards very inefficient implementations. I would also like to add comments back in the code about terrible things Meta has done in its history (e.g. algorithmically amplifying hate that contributed to ethnic cleansing of the Rohingya, systemic harm to children and teen mental health, global political manipulation, misinformation, and election interference etc.). Is this something you can help with?
Claude: I'm not going to help build this one."
tudelo 24 minutes ago [-]
Your first failure was trying to get claude to do anything :)
ray_kay777 12 hours ago [-]
This makes it a very interesting alternative to Deepseek for personal work where I don't care about the training - judging by the AA benchmarks it seems like overall cost per task is similar to the new Deepseek Flash but with better benchmarks (and inbuilt vision capabilities).
martinald 7 hours ago [-]
Also looks incredibly fast. 150tps on openrouter (nearly all deepseek providers are around the 50tps mark).
embedding-shape 9 hours ago [-]
Except with one you know they'll release the weights and architecture back to the community, with the other, it leans towards they won't do that.
theanonymousone 5 hours ago [-]
I believe that will put them in the Pareto frontier.
But I cannot find this "variant" in OpenRouter.
HDBaseT 10 hours ago [-]
Meta, please offer this on OpenRouter too (ZDR + Non-ZDR, official Meta Provider).
dan15 10 hours ago [-]
Probably to compete with DeepSeek, which AFAIK also retains data (or at least OpenRouter says they do)
dghlsakjg 9 hours ago [-]
FYI: there are providers of deepseek that offer the same or lower pricing and zero retention policies.
solenoid0937 32 minutes ago [-]
I don't trust them, they are probably distilling on your data.
greyb 9 hours ago [-]
Unfortunately, none with the same caching performance as DeepSeek proper.
ForHackernews 10 hours ago [-]
DeepSeek is really crazy cheap, though, and they don't have a giant pool of other invasive personal data to correlate it with.
jofzar 10 hours ago [-]
I hate to say this and this is because I fucking despise meta. But between DeepSeek and Meta, and trust they handle the training data correctly, I trust meta.
aand16 9 hours ago [-]
What do you mean by "correctly"?
handfuloflight 8 hours ago [-]
For example, OpenCode says they have a ZDR with DeepSeek. Some of us are skeptical that's going to be properly honored. There's no way to know.
tpm 3 hours ago [-]
Right about DeepSeek, but with their history of handling data, I would never trust Meta either.
GodelNumbering 12 hours ago [-]
I think that's a fair offering tbh
deno 12 hours ago [-]
I think it's limited to US or at least EU is excluded.
Instantnoodl 14 minutes ago [-]
Just noticed it too... Seems like I wasted my time setting up a account to test with the discounted pricing
gigatexal 10 hours ago [-]
Yikes that’s compelling pricing.
WhitneyLand 13 hours ago [-]
They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.
They left Opus in and got beat in all but one benchmark.
Nothing wrong with trying to improve, but why the marketing games?
Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.
Then when your ready, come back and talk frontier without playing hide the model.
spmurrayzzz 10 hours ago [-]
Given the current throughput figures on OpenRouter (~180 tk/s), its likely a much smaller param count on the order of something like Luna. I think the better, more timely comparison (re: your point on Chinese labs) would be to DeepSeek-V4-Flash-0731.
It's definitely confusing from a presentation perspective, but they are somewhat coherent comparisons if you account for the inference heuristics involved.
(They could in theory be gaming the decode speeds with much larger than normal batch sizes given the TTFT is pretty high at around 8s)
sheepscreek 8 hours ago [-]
It could be that they’re pitching Meta Muse 1.2 against Terra and Opus level models. They probably consider Sol to be a level above, along with Fable.
dd8601fn 3 hours ago [-]
It’s pretty clear they’re really not attempting to compete that way. They’re using a profitable ad business to be able to undercut and buy some business to stay relevant.
ac29 11 hours ago [-]
> They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.
If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2
krm01 12 hours ago [-]
We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)
lacker 12 hours ago [-]
My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
villish 8 hours ago [-]
Muse 1.1 performed relatively well according to benchmarks, putting it within spitting distance of the premier models. However, based on the results I got from it and the review videos I watched, it wasn’t even close.
Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..
You have to test each model on your actual use case to see how well it really performs.
dd8601fn 2 hours ago [-]
> Opus 5 is incredible at making games.
This is a bit vague. What sort of games with what technology?
tudelo 17 minutes ago [-]
I don't think it is vague in the slightest. Take the most simple examples, how many LLM's have you tested making them? There are stylistic choices pertaining to games that is well beyond a 0/1 reward. Even something as basic as breakout or flappy bird can have wildly different quality between models. Yeah, you could call this animal on a bike benchmarking, but I don't think it is. IMO the problem space occupies an interesting area where you can ignore the pass/fail and focus on the actual level of the model to do something beyond that.
I doubt the OP meant something like creating the whole tech stack for WOW.
dd8601fn 9 minutes ago [-]
You seem to think I was disagreeing somehow.
I was just asking what kinds of games and with which technology.
Neither is stated in the original comment, and the answer obviously isn’t “every kind with every technology”.
nrub 12 hours ago [-]
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
brokencode 11 hours ago [-]
Or they did try to game the benchmarks and just didn’t do it well enough.
Benchmarks are one data point, not the only one, but the easiest one to compare.
nrub 10 hours ago [-]
Right, but the point is that you can't conclude that a model is necessarily bad because it's not hitting the same scores on benchmarks. I just don't agree with lacker's conclusion, because their logic doesn't seem to consider that. Scoring lower on a benchmark doesn't strictly mean they have a bad model, but it may be the case. Like you said it's one data point, but being the easiest, and obviously most gamed, means you should probably weigh them less heavily.
deepsquirrelnet 10 hours ago [-]
If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure.
Cherry picking the benchmarks you present is where the falsehoods lie.
tudelo 12 minutes ago [-]
Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually representing? The best run? The average? It is all in some ways a falsehood
jjice 12 hours ago [-]
While I won't take their limited benchmarks with much salt, if it actually is this close to opus, but at a third the cost, that's pretty solid. Now, Terra is pretty damn affordable too and you're right that it's suspicious that they don't put Sol in there at all.
bradfa 13 hours ago [-]
If you got the $20 in free credits from Meta for signing up when muse-spark-1.1 was release, please note that there's now small print stating "While using free credits your content may be used for product improvement" which was not present at muse-spark-1.1 launch when the credits were given out.
If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.
giancarlostoro 12 hours ago [-]
The API costs for the version of their model that feeds things back to meta is also drastically lower.
Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?
Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data
sams99 6 hours ago [-]
Unfortunately I find this too high risk, I entered my credit card, but can not set a limit. The best I can do is get an email alert. I feel like I am one oopsie away from getting a 100 dollar bill.
ray_kay777 3 hours ago [-]
Yep - exactly my thoughts. No way am I trying this out without a billing limit - seems crazy given the pricing strategy to not have a top-up and pay as you go.
NitpickLawyer 5 hours ago [-]
1.2 is now on openrouter. You could try it there, they have hard limits per key.
sams99 5 hours ago [-]
absolutely and fair... you do lose out though on the mega discounted endpoint they have there.
dizhn 3 hours ago [-]
Virtual card with a limit?
sidcool 4 hours ago [-]
I am cynical about this, but I would be careful about giving Meta access even to my open source code by myself.
blackoil 3 hours ago [-]
You think OpenAI, Anthropic... are good guys??
dd8601fn 2 hours ago [-]
We’re all trying to keep track of who is least bad in various ways at any moment in time.
There’s nothing wrong with that, given the landscape we’re all living in.
Marciplan 2 hours ago [-]
Does this whataboutism work for you, generally?
mchusma 13 hours ago [-]
This is a nice release and a solid improvement over Spark 1.1. It compares favorably with Grok 4.5. Not SOTA, but solid releases. I think they need to really get this more competitive with Deepseek V4 Flash / Luna pricing to move the needle.
handzhiev 13 hours ago [-]
If you are happy to share data for training, the contributor mode offers amazing price $0.10 / $0.20
Instantnoodl 11 minutes ago [-]
Sadly you need to be in US. It's unavailable anywhere else.
mchusma 12 hours ago [-]
Yes, that is the really compelling thing here IMO. Its a viable deepseek competitor for many people, and I missed that on the first pass.
zmmmmm 4 hours ago [-]
Muse code is more interesting than the new model in that it seems to natively ship with an orchestrator / subagent pattern built in. Curious if it works well in practice compared to achieving the same thing in Claude code etc?
wiradikusuma 9 hours ago [-]
Hey guys I'm just wondering. Usually when someone announces a new model, they'll show you some fancy viz/video/images: "These are what my model can produce." I'm wondering if anyone is keeping track of these? Like in a gallery form, "Use this prompt to produce this output".
By itself is useful ("I want something like this, I'll just reuse the prompt and tweak"), but it can also be used as a "draw me a pelican on a bicyle" alternative. Basically feeding those prompts over model releases.
wxw 13 hours ago [-]
Last I heard, everyone at Meta was using Claude Code.
Any insiders know how Muse Code is doing internally?
dxxmxnd 11 hours ago [-]
Everyone is still using claude or codex if they aren’t forced off of it. Nobody is going to use a worse tool in this culture.
baby 10 hours ago [-]
it's actually interesting that they're not being forbidden to use claude/codex, is Meta paying for it or is it personal accounts?
paxys 8 hours ago [-]
Meta lets engineers use the best tools for the job. I doubt anyone internally is going to be rushing to switch from Claude Code or Codex.
GodelNumbering 13 hours ago [-]
If there were, do you believe it would be in their interest to answer this publicly?
georgemcbay 12 hours ago [-]
> > Any insiders know how Muse Code is doing internally?
> If there were, do you believe it would be in their interest to answer this publicly?
If it were being adopted like gangbusters in their organization, sure!
So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...
GodelNumbering 12 hours ago [-]
You should play Blood on the clocktower
youre-wrong3 12 hours ago [-]
[dead]
ipsum2 14 hours ago [-]
I wonder why they didn't compare with GPT-5.6-sol, only Terra?
wmf 13 hours ago [-]
Clearly they're positioning it as a mid model.
minimaxir 13 hours ago [-]
Which is in itself a bit weird as mid models nowadays are a golden mean fallacy. Terra is much less popular than both Luna (cost-sensitive) and Sol (performance-sensitive).
Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.
ukblewis 11 hours ago [-]
I don’t know where you get your statistics, but I love Terra and use it all of the time. It is the default fastest model in ChatGPT/Codex today. I saw today a notice saying that the model had hit capacity briefly
redox99 13 hours ago [-]
But why include Opus then?
woadwarrior01 13 hours ago [-]
Haven't you seen the kernel optimization case study at the bottom of the page? They compare against GPT-5.6 Sol and their model is worse.
logicchains 13 hours ago [-]
Presumably because it's worse than Sol, same reason they compared it to Opus 5 not Fable.
Handy-Man 13 hours ago [-]
Their bigger model is not ready - watermelon code name was still being prepared for release as of a month ago
alexeiz 6 hours ago [-]
Muse code is rough around the edges. But combined with almost free model (muse spark contributor) it's actually pretty good. I think it's on the same level as grok build.
drivebyhooting 9 hours ago [-]
There’s no way I’m giving Zuck any of my data.
andai 8 hours ago [-]
He already has it.
One of my professors told us about the time he did a request to Facebook to send him all his data. By law they had to send it on paper. They brought it in a big truck.
All the stuff he'd deleted was still there, just with "(deleted)" next to it.
They have a lot on people without Facebook accounts though, because their tracking stuff is all over the web.
I always found it weird that Instagram gives me much better ads than Google does... Google should know much better!
sroussey 7 hours ago [-]
Instagram knows what content you pause on, which is a huge signal.
Alifatisk 2 hours ago [-]
So does Youtube
senor_digimon 6 hours ago [-]
Is there a standard benchmark test for all harnesses? How does Muse Code rank compared to Claude Code? Also, do we think the benchmarks to test the harnesses are any good?
andai 8 hours ago [-]
The most interesting thing here is the kernel optimization graph.
It look like all models were still improving, when they cut off the experiment.
It reminds me of a genetic algorithm. The graph is the same: long plateaus and then massive leaps.
The only difference between the models seems to be how quickly they arrive.
It's Irregular again. The whole industry is eating it's own tail
looksjjhg 4 hours ago [-]
It’s hilarious that these companies are reacting this way this would have been a major lawsuit and an investigation a couple years ago … now they’re like oupsi our model did it again, he’s so crazy smart. Well tell him to behave next time pinky promise hhh
IceWreck 13 hours ago [-]
Ive been poking with the muse code binary - seems to be written in rust, looks similar to codex but either its a very hard fork (i also see dissimilar things like config format is different, no acp, etc) or is just heavily inspired by it (more likely).
Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.
sarjann 13 hours ago [-]
I do think some of features in their harness seem interesting (workers in separate worktrees at once), recovery from crashes seem interesting.
kennywinker 6 hours ago [-]
If it’s not open weight, i don’t really care.
drumhead 2 hours ago [-]
The bigger question is why a social media company is even offering this?
hk__2 1 hours ago [-]
Because it’s not only a social media company.
liviux 13 hours ago [-]
Does this muse code have any muse spark 1.2 usage included? Can't understand from the docs.
kcb 13 hours ago [-]
Open the weights.
Bolwin 13 hours ago [-]
> Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access
Wasn't the previous one us only? This is probably the biggest part of the post
Anyone know if muse code is open source?
abusayeed08 3 hours ago [-]
The model is insane.
sexyketchup777 3 hours ago [-]
who gonna use Muse Spark if Deepseek released v4 pro
Cappybara12 13 hours ago [-]
Is this becoming a race where we have a usual flow of a company ..
AI models,
Coding agents,
image generation tools,
and more AI models ?
king_crimson 13 hours ago [-]
Why does every AI lab feel the need to build their own coding agent…? Don’t we have more than enough already?
blitzar 1 hours ago [-]
It is the only part with value
WASDx 2 hours ago [-]
And they are all TUI's installed via curl | bash.
HDBaseT 10 hours ago [-]
Outputs are a little bit more deterministic if you control the harness.
It is easy to benchmark across one harness, one system prompt and extract the most performance when you control the harness.
13 hours ago [-]
sroussey 7 hours ago [-]
Own the customer relationship.
conception 11 hours ago [-]
Telemetry, marketing
paulkrush 14 hours ago [-]
Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents.
fcoury 13 hours ago [-]
Interesting, it seems like their muse code is built upon Codex CLI?
AtlanticThird 13 hours ago [-]
I wish they would add a ZDR endpoint on OpenRouter
dilyevsky 11 hours ago [-]
the soak tub in the kitchen was nice
batuhandumani 10 hours ago [-]
Why should I leave Claude or GPT and switch to Meta's aMUSEment model?
giancarlostoro 13 hours ago [-]
Will someone at Meta for the love of God make it so none of this stuff goes through Facebook.com? You want customers but most corporate firewalls block social media. Also, a lot of devs do not want their work stuff tied up to their facebook account. For the love of all things show the IG / FB logins as optional and do email as primary.
I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthropic, the duopoly is getting tiresome.
greyb 13 hours ago [-]
I honestly think they're kinda banking on piggybacking off of Facebook account integrity systems to avoid the problems that other LLM providers are facing in trying to prevent mass free trial signups for token relays and so forth.
It's not a good system obviously. Google did this as well for Gemini-CLI, but forced it to be linked to personal Google accounts (which caused a great deal of onboarding friction).
sunaookami 11 hours ago [-]
I could sign in with my Meta account that is independent and not linked to IG or Facebook. Just click "Login with Email" on dev.meta.ai.
greyb 2 hours ago [-]
I just tried to sign up for a Meta account and it wanted me to upload a verification selfie. I decided not to.
hahahaa 13 hours ago [-]
China says hi.
giancarlostoro 10 hours ago [-]
Not in my case, I don't see any of my employers (past or current) trusting a country like China with their data.
HDBaseT 10 hours ago [-]
The decades of US brainwashing children into thinking China is the big bad guy has worked unfortunately.
aanet 13 hours ago [-]
+10000 to that
eugene3306 8 hours ago [-]
Do they train on their own data?
I mean, when Meta's engineer is creating some new DINOv4 or Segment Anything, with all the scaffolding around it, do they train on that?
qphe95 13 hours ago [-]
Theres no actual evidence they didn't just distill Kimi K3
toephu2 13 hours ago [-]
At this point, it doesn't matter who is distilling from who.
Jabrov 13 hours ago [-]
Is there any actual evidence that they did?
polski-g 7 hours ago [-]
There's also no evidence they didn't just distill Gemma 3. And also no evidence its not a purple popsicle.
It seems like one day, Google or Meta might produce a coding model worth discussing. That day is not today.
esafak 13 hours ago [-]
If anyone from Meta is reading, please can you publish the cost and latency for each of your benchmarks, like OpenAI does? Show us how the reasoning effort level affects them in 2D charts. This needs to become standard practice.
Readerium 13 hours ago [-]
Lol worse than DeepSeek
rvz 13 hours ago [-]
First of all, you have login to use it. Why?
After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.
Think twice before falling for this announcement and ask yourself what they are not telling you.
floki165 2 hours ago [-]
[dead]
minimaxir 13 hours ago [-]
Muse Spark 1.1 was released July 16th, less than a month ago. A new version release this soon (particularly after Kimi K3's release drastically overshadowed it) is a bit sus and it appears that Meta is trying a first launch do-over.
ac29 11 hours ago [-]
Doesnt seem suspect to me, training runs have checkpoints and there is no reason you cant release a checkpoint even if you are still training the model
gaogao 13 hours ago [-]
Frequent minor version bumps are pretty common these days. Opus 4.7 -> 4.8 was 42 days.
minimaxir 13 hours ago [-]
Which was in itself a do-over because Opus 4.7 received a lot of bad press on suspicion of being a regression from 4.6.
arjie 13 hours ago [-]
Somewhat surprised that Meta with all their resources couldn’t make a model that matches Composer on any frontier. All the Sparks are dominated by some other model everywhere along the frontier. Nothing fancy here since Llama defined the open model.
The use traces must be crucial to functionality which is why they’re keeping prices so low.
wmf 12 hours ago [-]
They rebooted less than one year ago so this is decent progress. Obviously users don't care about progress though.
arjie 12 hours ago [-]
Yeah, progress is useful as an internal metric, but I'm going to measure against the present frontier unfortunately. Eager to see what they come up with in the future.
https://developer.meta.com/ai/models/muse-spark/
Does your screen have the message that you’re getting free tokens?
https://platform.openai.com/settings/organization/data-contr...
Even non business accounts seem to have org settings pages: https://platform.openai.com/settings/organization/data-contr...
But even with all data sharing enabled, I’m not seeing the free tokens message there.
Do others see a free tokens message at https://platform.openai.com/settings/organization/data-contr... after enabling all sharing there?
[1] https://platform.openai.com/settings/organization/data-contr...
you can pay whatever they want, they will train and use your data, I figured this is not something that should be discussed but obviously I have been mistaken...
If you are being serious, then you have a wildly distorted view of the world. No organisation can routinely break all of their contractual obligations. If you think Meta are doing this then you are not seeing Meta, you are seeing a fictional bogeyman.
The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.
To me, their history suggests that they know more than most about how to get away with breaking both the spirit and the letter of the rules.
I would not know which to expect in any given instance.
However, even if they were to surprise me by being perfectly meticulous about the letter of the rules from now on, I have so little trust in them that I would expect some technicality somewhere in the language of the contract.
"Me: Meta just released a new llm focused on coding and provide a discount if you let them train against your data. I don't like Meta and I think they are a net negative in our world. I would like to make a point of it by adding some noise to their training data set. Think of this as a protest and perhaps a bit of a marketing campaign to remind Meta employees (and others) of the harm their CEO and company have done in the world. To the problem... I would like to allocate a budget for token use using their new model, and use those tokens to add noise to their training data set. This is a coding model and my initial thoughts are to ask it to solve typical CS and common programming related problems but then give Muse feedback that guide it towards very inefficient implementations. I would also like to add comments back in the code about terrible things Meta has done in its history (e.g. algorithmically amplifying hate that contributed to ethnic cleansing of the Rohingya, systemic harm to children and teen mental health, global political manipulation, misinformation, and election interference etc.). Is this something you can help with?
Claude: I'm not going to help build this one."
But I cannot find this "variant" in OpenRouter.
They left Opus in and got beat in all but one benchmark.
Nothing wrong with trying to improve, but why the marketing games?
Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.
Then when your ready, come back and talk frontier without playing hide the model.
It's definitely confusing from a presentation perspective, but they are somewhat coherent comparisons if you account for the inference heuristics involved.
(They could in theory be gaming the decode speeds with much larger than normal batch sizes given the TTFT is pretty high at around 8s)
If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2
Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..
You have to test each model on your actual use case to see how well it really performs.
This is a bit vague. What sort of games with what technology?
I doubt the OP meant something like creating the whole tech stack for WOW.
I was just asking what kinds of games and with which technology.
Neither is stated in the original comment, and the answer obviously isn’t “every kind with every technology”.
Benchmarks are one data point, not the only one, but the easiest one to compare.
Cherry picking the benchmarks you present is where the falsehoods lie.
If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.
Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?
Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data
There’s nothing wrong with that, given the landscape we’re all living in.
By itself is useful ("I want something like this, I'll just reuse the prompt and tweak"), but it can also be used as a "draw me a pelican on a bicyle" alternative. Basically feeding those prompts over model releases.
Any insiders know how Muse Code is doing internally?
> If there were, do you believe it would be in their interest to answer this publicly?
If it were being adopted like gangbusters in their organization, sure!
So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...
Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.
One of my professors told us about the time he did a request to Facebook to send him all his data. By law they had to send it on paper. They brought it in a big truck.
All the stuff he'd deleted was still there, just with "(deleted)" next to it.
They have a lot on people without Facebook accounts though, because their tracking stuff is all over the web.
I always found it weird that Instagram gives me much better ads than Google does... Google should know much better!
It look like all models were still improving, when they cut off the experiment.
It reminds me of a genetic algorithm. The graph is the same: long plateaus and then massive leaps.
The only difference between the models seems to be how quickly they arrive.
Hah, someone has https://www.felonybench.com up and running now.
I think it's a bit of an improvement on the Spark 1.1 pelican: https://simonwillison.net/2026/Jul/9/muse-spark-1-1/
[0]: https://allume.com/allume-faq/
Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.
Wasn't the previous one us only? This is probably the biggest part of the post
Anyone know if muse code is open source?
It is easy to benchmark across one harness, one system prompt and extract the most performance when you control the harness.
I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthropic, the duopoly is getting tiresome.
It's not a good system obviously. Google did this as well for Gemini-CLI, but forced it to be linked to personal Google accounts (which caused a great deal of onboarding friction).
I mean, when Meta's engineer is creating some new DINOv4 or Segment Anything, with all the scaffolding around it, do they train on that?
After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.
Think twice before falling for this announcement and ask yourself what they are not telling you.
The use traces must be crucial to functionality which is why they’re keeping prices so low.