Hi everyone,
I have been comparing AI girlfriend apps, and nearly all of them now promote “long-term memory.” However, a larger context window does not always result in better recall or more consistent conversations.
Common problems include forgetting preferences, contradicting earlier details, repeating responses and unexpected personality changes.
A practical benchmark could test:
- Recall after 20–50 messages
- Memory following a topic change
- Correction of inaccurate information
- Consistency across multiple sessions
- The ability to delete stored details
For developers, is native context enough, or do these applications need summaries, retrieval systems and structured user profiles?
Are there any NVIDIA NeMo or NIM evaluation methods for measuring memory failures and conversational consistency?
I’d be interested to hear how others evaluate memory and conversation consistency in these applications.