This kind of “introspection” with LLMs is useless. It’s always just predicting the next token, even when asked about why it did something. The response doesn’t need to be at all related to why it did that. It hasn’t been trained to analyze how it predicts each token, though it was trained on a bunch of other recorded conversations of people being scolded and questioned about getting something wrong. It’s just spitting out a version of that when you get upset with it. Similar for talking about how it works, it’s likely regurgitating conversations about how LLMs function, though it could get caught in a context where someone is talking about how they (a human) function.
If you’re doing it for any reason other than your own amusement, you’re just wasting time and filling the context window with junk.
This kind of “introspection” with LLMs is useless. It’s always just predicting the next token, even when asked about why it did something. The response doesn’t need to be at all related to why it did that. It hasn’t been trained to analyze how it predicts each token, though it was trained on a bunch of other recorded conversations of people being scolded and questioned about getting something wrong. It’s just spitting out a version of that when you get upset with it. Similar for talking about how it works, it’s likely regurgitating conversations about how LLMs function, though it could get caught in a context where someone is talking about how they (a human) function.
If you’re doing it for any reason other than your own amusement, you’re just wasting time and filling the context window with junk.