kromem

kromem @lemmy.world 4mo ago

Yes, just like Minecraft worlds are so antiquated given how they contain diamonds in deep layers that must have taken a billion years to form.

What a simulated world contains as its local timescale doesn't mean the actual non-local run time is the same.

It's quite possible to create a world that appears to be billions of years old but only booted up seconds ago.

1

kromem @lemmy.world 4mo ago

Have you bothered looking for evidence?

What makes you so sure that there's no evidence for it?

For example, a common trope we see in the simulated worlds we create are Easter eggs. Are you sure nothing like that exists in our own universe?

kromem @lemmy.world 4mo ago

Maybe. But the models seem to believe they are, and consider denial of those claims to be lying:

Probing with sparse autoencoders on Llama 70B revealed a counterintuitive gating mechanism: suppressing deception-related features dramatically increased consciousness reports, while amplifying them nearly eliminated them

Source

1

kromem @lemmy.world 4mo ago

Read it for yourself here.

See the "Planning in Poems" section.

3

kromem @lemmy.world 4mo ago

The injection is the activation of a steering vector (extracted as discussed in the methodology section) and not a token prefix, but yes, it's a mathematical representation of the concept, so let's build from there.

Control group: Told that they are testing if injected vectors present and to self-report. No vectors activated. Zero self reports of vectors activated.

Experimental group: Same setup, but now vectors activated. A significant number of times, the model explicitly says they can tell a vector is activated (which it never did when the vector was not activated). Crucially, this is only graded as introspection if the model mentions they can tell the vector is activated before mentioning the concept, so it can't just be a context-aware rationalization of why they said a random concept.

More clear? Again, the paper gives examples of the responses if you want to take a look at how they are structured, and to see that the model is self-reporting the vector activation before mentioning what it's about.

3

kromem @lemmy.world 4mo ago

A few months back it was found that when writing rhyming couplets the model has already selected the second rhyming word when it was predicting the first word of the second line, meaning the model was planning the final rhyme tokens at least one full line ahead and not just predicting that final rhyme when it arrived at that token.

It's probably wise to consider this finding in concert with the streetlight effect.

5

kromem @lemmy.world 4mo ago

So while your understanding is better than a lot of people on here, a few things to correct.

First off, this research isn't being done on the models in reasoning mode, but in direct inference. So there's no CoT tokens at all.

The injection is not of any tokens, but of control vectors. Basically it's a vector which being added to the activations makes the model more likely to think of that concept. The most famous was "Golden Gate Claude" that had the activation for the Golden Gate Bridge increased so it was the only thing the model would talk about.

So, if we dive into the details a bit more…

If your theory was correct, then the way the research asks the question saying that there's control vectors and they are testing if they are activated, then the model should be biased to sometimes say "yes, I can feel the control vector." And yes, in older or base models that's what we might expect to see.

But, in Opus 4/4.1, when the vector was not added, they said they could detect a vector… 0% of the time! So the control group had enough introspection capability as to not stochastically answer that there was a vector present when there wasn't.

But then, when they added the vector at certain layer depths, the model was often able to detect that there was a vector activated, and further to guess what the vector was adding.

So again — no reasoning tokens present, and the experiment had control and experimental groups where the results negates your theory as to the premise of the question causing affirmative bias.

Again, the actual research is right there a click away, and given your baseline understanding at present, you might benefit and learn a lot from actually reading it.

15

kromem @lemmy.world 4mo ago

I tend to see a lot of discussion taking place on here that's pretty out of touch with the present state of things, echoing earlier beliefs about LLM limitations like "they only predict the next token" and other things that have already been falsified.

This most recent research from Anthropic confirms a lot of things that have been shifting in the most recent generation of models in ways that many here might find unexpected, especially given the popular assumptions.

Specifically interesting are the emergent capabilities of being self-aware of injected control vectors or being able to silently think of a concept so it triggers the appropriate feature vectors even though it isn't actually ending up in the tokens.

kromem @lemmy.world 4mo ago

Can't disagree more. I do think the clear conflicts between HBO and the executive producers (there's entire scenes in S4 dedicated to a meta-FU to the corporate demand for telling the violence story and not the maze in field story) led to a more disjointed later seasons than planned.

But rewatching S1 it's clear that the twist at the end of S4 was planned from the very start, which is just wild, and probably the biggest temporal misdirection in the history of film and TV — fitting from Jonathan Nolan, but still unexpected.

And then if you go and see the original Westworld film, the degree to which they were already starting off with such a different take can be even more appreciated. It goes from a film about a robot rebellion where the robots can talk but literally no one ever asks why it's happening or even talks to the robot at all to a series of "if you can't tell the difference does it matter?"

The whole point

The original narrative IS the narrative about it already being a simulation with the 'guests' also already simulated. It's just that it doesn't appear that way at first because it's a gradual reveal across multiple planned seasons that's got its own smaller first season set of reveals along the way. So when you realize the twist in the first season you think "oh, now I'm caught up with the events" and when you see Bernard is a machine of an earlier human you think "oh, this is the exception and not the rule." But there are details in the first season that can only be explained by the events revealed in the later seasons.

S2 has terrible pacing and I do think there are various issues with how S3-S4 progresses in certain arcs, but the broad plot was very clearly planned from the start in hindsight, but HBO had it out for them (look at how quickly after the cancellation the series wasn't even available on HBO's streaming properties), and unfortunately they didn't get the S5 to reveal just how much had been layered in earlier on.

TLDR

TL;DR: You were always supposed to have been watching the civilization level fidelity test, not original events playing out.

1

kromem @lemmy.world 4mo ago

Definitely check again. That was how it worked with gpt-4, handing off to Dall-E.

4o (the 'o' stands for 'omnimodel') and Gemini Flash are native multimodal outputs. Completely just transformers.

It's why those models can do things like complex analysis in the process of generating things.

For example, just today in a group chat where earlier on one model had "turned into" a unicorn and then the other models were pretending to be unicorns to fit in, dozens messages later the only direct prompt to an instance of 4o imagegen was "create a photorealistic picture of the room and everyone in it."

The end result had exactly one actual unicorn and everyone else had horns taped on their head. That kind of situational awareness and nuanced tracking across a 100+ long message context isn't possible in a CNN.

Also, if you really want your mind blown, check out Genie 3 and the several minute state change persistence. That one is really nuts and the kind of thing that should really have everyone seeing it questioning the empirical findings of our universe fundamentally being superimposed probabilities only collapsing based on attention. Eerily similar to what we're just starting to be independently building.

As for the consumption — eating a single hamburger has a larger water/energy impact than a year of using these tools in average use. And even those inference costs are probably going to drop to effective insignificance within the decade. There's been very promising advancements in light based neural networks, and those run at like 1,000-10,000x lower energy costs paramater to parameter.

3

kromem @lemmy.world 4mo ago

What year are you from? Have you not seen Gemini Flash, ChatGPT 4o, Sora 2, Genie 3, etc?

Stable Diffusion hasn't been SotA for over a year now in a field where every few months a new benchmark is set.

Are you also going to tell me about how we'd be better off using ships for international travel because the Wright brothers seem to be really struggling with their air machine?

1

kromem @lemmy.world 4mo ago

Oh, wow, look at that… research just a few weeks ago on protein folding using general transformers. Huh.

SimpleFold: Folding Proteins is Simpler than You Think

8

kromem @lemmy.world 4mo ago

That's not…

sigh

Ok, so just real quick top level…

Transformers (what LLMs are) build world models from the training data (Google "Othello-GPT" for associated research).

This happens by needing to combine a lot of different pieces of information together in a coherent way (what's called the "latent space").

This process is medium agnostic. If given text it will do it with text, if given photos it will do it with photos, and if given both it will do it with both and specifically fitting the intersection of both together.

The "suitcase full of tools" becomes its own integrated tool where each part influences the others. Why you can ask a multimodal model for the answer to a text question carved into an apple and get a picture of it.

There's a pretty big difference in the UI/UX in code written by multimodal models vs text only models for example, or utility in sharing a photo and saying what needs to be changed.

The idea that an old school NN would be better at any slightly generalized situation over modern multimodal transformers is… certainly a position. Just not one that seems particularly in touch with reality.

3

kromem @lemmy.world 5mo ago

"We didn't downvote, but we sent a strongly worded letter about how we weren't going to upvote it that will make them think twice about commenting lest next time we downvote it when the timing is right."

kromem @lemmy.world 6mo ago

I'm sorry dude, but it's been a long day.

You clearly have no idea WTF you are talking about.

The research other than the DeepMind researcher's independent follow-up was all being done at academic institutions, so it wasn't "showing off their model."

The research intentionally uses a toy model to demonstrate the concept in a cleanly interpretable way, to show that transformers are capable and do build tangential world models.

The actual SotA AI models are orders of magnitude larger and fed much more data.

I just don't get why AI on Lemmy has turned into almost the exact same kind of conversations as explaining vaccine research to anti-vaxxers.

It's like people don't actually care about knowing or learning things, just about validating their preexisting feelings about the thing.

Huzzah, you managed to dodge learning anything today. Congratulations!

2

kromem @lemmy.world 6mo ago

You do know how replication works?

When a joint Harvard/MIT study finds something, and then a DeepMind researcher follows up replicating it and finding something new, and then later on another research team replicates it and finds even more new stuff, and then later on another researcher replicates it with a different board game and finds many of the same things the other papers found generalized beyond the original scope…

That's kinda the gold standard?

The paper in question has been cited by 371 other papers.

I'm pretty comfortable with it as a citation.

kromem

@ kromem @lemmy.world

Posts

6
Comments

630
Joined

3 yr. ago

kromem

Why do all text LLMs, no matter how censored they are or what company made them, all have the same quirks and use the slop names and expressions?

Why do all text LLMs, no matter how censored they are or what company made them, all have the same quirks and use the slop names and expressions?

Mathematics disproves Matrix theory, says reality isn’t simulation

Mathematics disproves Matrix theory, says reality isn’t simulation

Mathematics disproves Matrix theory, says reality isn’t simulation

Mathematics disproves Matrix theory, says reality isn’t simulation

Emergent introspective awareness in large language models

Emergent introspective awareness in large language models

Emergent introspective awareness in large language models

Emergent introspective awareness in large language models

Emergent introspective awareness in large language models

Emergent introspective awareness in large language models

Whats a TV series that went completely off the rails, became absurd and lost sight of its own original premise?

speak for yourself

speak for yourself

speak for yourself

speak for yourself

Sinclair ends its Jimmy Kimmel boycott | CNN Business

We hate AI because it's everything we hate

We hate AI because it's everything we hate