The Need to Understand the Next Step


Rudina Seseri

Three months ago, I discussed how Genie 3’s developments were advancing the world model ecosystem. One prevailing issue that remains with world models is their inclination to hallucinate. Frequently, these models are not trained in specific environments or not on edge cases within them. The models respond by completing a task confidently wrong and not alerting the user of a potential issue. A recent paper from a team at the University of California San Diego proposes that hallucinations are a data coverage issue, and hallucinations can be caught and addressed earlier, making world model training faster and cheaper.

Given that world models are an AI architecture that builds an internal representation of their surroundings to predict outcomes and decision-making, the assumption has been that making world models more accurate required larger architectures and more training compute. The UCSD researchers posit that the issue is a data coverage issue, where there is not enough specific situational data to complete a task correctly. Solving for specific gaps is a much quicker and simpler task for researchers and engineers.

🗺️ What is data coverage?

Data coverage refers to a data set containing both the situation and choice a model will encounter. An issue with world models today is that they hallucinate answers, like a tour guide describing an area of a museum they have never visited. Since world models undertake multiple internal stages of processing, the error appears as different engineering issues to solve, rather than as a single problem.
The UCSD researchers determined that the data coverage issue can be identified by looking at how consistent a model is with itself across three metrics:

  • Round trip residual: This metric resembles asking a model to translate a sentence given in English to French and back to English again to see if the same starting point is achieved. World models are performing this task with predicted images and internal shorthand.
  • Flow instability: The second test looks at the confidence of a model as it is generating the next frame. Each time a world model generates the next frame in a sequence it is internally taking multiple passes. If the image changes drastically every time in the final passes, it is clearly a lack of confidence in the answer.
  • Inter-seed variance: This last metric tests whether a model can respond to the same question with the same answer several times over even when starting from a different point, like driving home but starting at different streets.

Testing a model along these three metrics will help the user determine where the model lacks confidence the most, since each of these metrics can be quantified. This process enables low confidence findings to be translated into gaps to address.

🤔 What is the significance of data coverage, and what are its limitations?

As world models continue to proliferate, addressing hallucinations quickly and cost effectively through data coverage will become increasingly valuable:

  • Targeted additions: Instead of needing to collect a large swath of new data, identifying data coverage gaps unlocks lower model and training costs because the incremented data is needed only where there is an identified gap.
  • Data set alignment: Given that these models are leveraging massive data sets, when training a model, the same or similar tasks can be randomly selected repeatedly, leading the model to not be trained on a subset of existing data. Auditing the data set reveals the skew, enabling engineers to resample the training data instead of collecting new data.
  • Better triaging: The initial quantified triaging provided by the model to steer automated data collection could also help engineers generate a prioritized list of data gaps to address.

While the ability to identify data coverage gaps is a meaningful step towards enabling world models to have fewer hallucinations, this method still has limitations:

  • Site specific context: New situations still require training. Being able to identify data coverage gaps still means they need to be filled, and every situation will have new gaps.
  • Diverse data requirements: Even though this targeted data coverage approach requires less data, large and diverse data sets are still needed, which are costly to create and store.
  • Human judgment: The paper notes that letting the model pick where to practice was better than random selection but fell short of a human discerning where best to train next.
🛠️ Use cases of data coverage

While this analysis only discusses a model that has yet to be tested in the physical world, the implications could have a great deal of impact, particularly in the robotics space:

  • Warehouse and fulfillment robotics: The data coverage approach enables world models, which are leveraged by robots, to be tested prior to their use to determine where gaps exist.
  • Autonomous driving: Today, autonomous driving companies collect large sets of data, but if the three metrics discussed were leveraged to test the model, smaller, more impactful data sets could be stored and leveraged.
  • Surgical robotics: A world model that can assess the confidence of an outcome is much more useful in medical settings, where a robot could be trained to flag for a human to step in.

Stay up-to-date on the latest AI news by subscribing to Rudina’s AI Atlas.