Semantic Scene Graphs for Creating a Localization-Ready Internet of Things

1 Visualization Research Center (VISUS), University of Stuttgart, Germany
2 Graz University of Technology, Austria
IEEE Transactions on Visualization and Computer Graphics (TVCG), vol. 32, no. 8, 2026
Presented at IEEE ISMAR 2026, the 25th International Symposium on Mixed and Augmented Reality, October 5–9, 2026, Bari, Italy
Conceptual illustration of IoT device localization and control in LORIOT.
Left: LLM-generated AR control panels (a) instantiated next to the physical devices. Right: the predicted associations (b) between network-side device identities and scene-graph object nodes.

Abstract

Controlling devices connected to the Internet of Things often requires juggling multiple smartphone apps or physical remote controls, creating a fragmented user experience. Augmented Reality (AR) can afford superior control by automatically presenting virtual user interfaces that are spatially aligned with networked devices. However, before such user interfaces can be delivered, physical devices must be localized in the environment. This paper introduces LORIOT (LOcalization-Ready Internet Of Things), a novel end-to-end system that uses a semantic scene graph and a large language model to map the identities of the networked devices to physical objects, given a pre-filtered set of IoT-device candidate nodes. A declarative UI specification enables automatic generation of device control panels for AR and non-AR clients. We evaluate the mapping component on a controlled synthetic-room benchmark of 100 randomly generated rooms. Using device network metadata alone, we achieve a baseline macro-averaged F1 score of 0.80 for digital → physical associations. When device metadata is enriched with physical attributes (mounting location, materials, color, and size), performance improves to 0.88. Moreover, we evaluate the benefit of spatially registered AR control in a within-subject user study (N=20), comparing in-situ AR panels against conventional non-AR control with smartphone apps or physical remote controls. AR yields significantly faster task completion, lower mental demand, and higher usability.

  • 42%faster task completion with AR39 s vs. 68 s per task with apps and remotes
  • 0.88macro F1, device → object mappingup from 0.80 with enriched device metadata
  • 84.3System Usability Scale with ARvs. 62.9 for smartphone apps and remotes
  • 20participants, 26 IoT deviceswithin-subject study in a real office

Motivation: IoT Control Is Fragmented

Smart environments fill up with connected lamps, fans, blinds and photo frames, but controlling them is still a fragmented experience. Every manufacturer ships its own smartphone app, some devices only come with a physical remote, and users end up guessing which app controls which lamp. Augmented Reality promises something better: a control panel that floats right next to the physical device. Before such an interface can be shown, however, the system has to answer two questions.

  • Where is it?The 6-DoF pose of every device in the room
  • Which one is it?The network identity behind each physical object

Establishing this link between a physical object and its network identity is what we call making the Internet of Things localization-ready.

Existing Methods Do Not Scale

Prior work bridges the gap between a device's network identity and its physical position through auxiliary channels, each with its own cost.

ApproachLimitation
Visual markers such as QR codes attached to every deviceMust be placed, kept visible and maintained
Radio signal strength (ZigBee, UWB)Imprecise; only works for a few devices that are far apart
Toggling device state and watching for the changeOnly for devices with visible state changes, not for buttons, blinds or height-adjustable desks

LORIOT relies on semantics instead: what the object looks like and what the device can do. Recent advances in open-vocabulary 3D scene understanding make this information available in a semantic scene graph, and a large language model can reason about it together with the device metadata found on the network.

LORIOT: An End-to-End Pipeline

  1. RGB-D scanof the environment
  2. Semantic scene graphobjects + attributes
  3. LLM device mappingdigital ↔ physical
  4. Generated AR UIsspatially registered
Network layout and client/server architecture of LORIOT.
Client/server architecture of LORIOT: the server creates the semantic scene graph, computes the device mapping, generates the user interfaces and routes all commands. AR and smartphone clients connect to it over the local network.

LORIOT runs as a distributed system over a local WiFi network. The server is built on Node-RED and an MQTT broker, with the device metadata stored in CouchDB. It constructs the semantic scene graph from an RGB-D scan of the room and uses a GPT-class large language model twice: first to map network identities to scene-graph objects, and second to generate user interfaces from device metadata. Two clients fetch these interfaces, an AR client running on a Meta Quest 3 and a conventional smartphone client written in Flutter. Every command a user issues is routed by the server to the devices over HTTP or MQTT.

From an RGB-D Scan to a Semantic Scene Graph

We build the scene graph with ConceptGraphs (Gu et al., ICRA 2024), which turns an RGB-D video into objects with 3D bounding boxes, open-vocabulary embeddings and detailed captions. LORIOT adds a list of descriptive tags and the dominant color to every node, which helps the language model to tell similar objects apart.

{
  "obj_id": 219,
  "object_tag": "lamp",
  "object_caption": "A small, modern black desk lamp with a conical shade, ...",
  "bbox_extent": [0.39, 0.17, 0.16],
  "bbox_center": [-1.95, 0.92, -0.26],
  "bbox_rot": [-0.75, 0.62, 0.22, -2.88],
  "object_tags": ["small", "modern", "black", "desk lamp",
                  "conical shade", "round base", "adjustable arm"],
  "color": [0.28, 0.25, 0.22]
}
One scene-graph node. The object tags and the color are added by LORIOT.
  • Manual refinement: duplicate detections are removed and wrong labels are fixed after the automatic step.
  • Pre-filtering: the IoT candidate nodes are selected by hand in our prototype. Automating this step is future work.

Device Mapping with an LLM

On one side is the set of metadata records of all networked devices, on the other side the set of scene-graph nodes that are IoT candidates. Ideally, the mapping between them is a bijection. We ask the LLM to rank the matching objects for each device (digital → physical) and then repeat the query in the opposite direction, ranking the matching devices for each object (physical → digital).

{
  "relevant_objects": [
    {"obj_id": 1111, "mapping_accuracy": 0.8},
    {"obj_id": 219,  "mapping_accuracy": 0.8}
  ],
  "query_achievable": true,
  "explanation": "Smart lamp. Objects 1111 and 219 are both
    small, modern desk lamps on desks; no other object matches."
}
LLM answer for one network device: ranked candidates with a confidence score, a flag whether the query is achievable at all, and an explanation that forces step-by-step reasoning.

Identical devices are the catch: if two identical lamps are present, no semantic description can tell which network identity belongs to which lamp. LORIOT therefore distinguishes instance-level associations, the two ranked directional relations with confidence scores, from a type-level bijection between device types and object types, which remains well defined. The remaining ambiguities are resolved in AR, where the user confirms or overrides a few ranked suggestions, far less work than labeling every device from scratch.

Automatic UI Generation

The same vendor-independent device metadata drives the UI generation. Each device is described by a JSON record with its name, network address and a list of commands, each with endpoint, request type and body. The LLM turns this record into a declarative UI description, and the AR client instantiates a Unity prefab for every widget. The pipeline is purely data-driven, so new devices need no code changes, and the smartphone client renders the very same JSON as a conventional 2D interface.

  • Boolean state change → toggle button
  • Query → button that shows the response text
  • Numeric parameter → slider with a $value placeholder in the command body
Generated control panel with annotations showing which widget issues which request.
A generated panel for a Nanoleaf light: toggle button (PUT request), state query (GET request with response text), brightness slider (GET request with parameters) and hue slider.
Automatically generated user interfaces for various IoT devices.
Generated control panels for a fan, motorized blinds, a push button, an RGB lamp, a Nanoleaf light panel and a digital photo frame, all produced by the same pipeline without any per-device code.

Registration and Voice Control

Two smaller pieces complete the system.

  • Spatial registration: a single fiducial marker aligns the coordinate system of the scan, and thus of the scene graph, with the headset tracking and relocalizes the AR client.
  • Natural-language control: for hands-free use, spoken requests such as “turn on the lamp near the window” are transcribed and interpreted by an agent that knows all device APIs, yielding a command the server can route directly.
  1. Spoken request“turn on the lamp near the window”
  2. Whisperspeech-to-text
  3. LangChain agentwith all device APIs from CouchDB
  4. Command JSONrouted by the server

Mapping Evaluation on 100 Synthetic Rooms

Rendering of a synthetic room with 20 labeled IoT devices.
One generated room with 20 devices.

Physical rooms do not scale for evaluation, so we built a room generator in Unity. It places devices of 11 types at plausible positions on walls, ceilings, floors and randomly sized desks, and copies their scene-graph descriptions from real scans of our office. Identical instances of a device type are collapsed into equivalence classes to obtain a type-level ground truth. We evaluated 100 rooms with 20 devices each and report precision, recall, F1 and IoU, macro- and micro-averaged, both unweighted and weighted by the confidence of the LLM.

  • 100randomly generated rooms
  • 20devices per room
  • 11device types
Type-level mapping quality, macro F1
Type-level mapping quality, macro F1
ItemNetwork metadata onlyEnriched metadata (mounting location, material, color, size)
Device → object (digital → physical)0.800.88
Object → device (physical → digital)0.810.93
Type-level bijection0.820.91
Weak device types benefit most, macro F1
Weak device types benefit most, macro F1
ItemNetwork metadata onlyEnriched metadata (mounting location, material, color, size)
Lamp0.420.85
Tasmota lamp0.470.65
Ceiling panel0.730.81
All scores are type-level: identical instances of a device type are collapsed into one class. 100 synthetic rooms × 20 devices.

With network metadata alone, mapping from device to object reaches a macro F1 of 0.80 and the reverse direction 0.81; the type-level bijection between both sides reaches 0.82. Enriching the metadata with mounting location, material, color and size, information a manufacturer could easily provide, raises these scores to 0.88, 0.93 and 0.91. The gains are largest for previously difficult types: the plain lamp improves from 0.42 to 0.85.

User Study: AR vs. Conventional Control

3D scan of the office with the positions of all 26 devices marked.
Our office with all 26 device placements.

Does knowing device locations actually help users? We ran a within-subject study with 20 participants (15 male, 5 female, mean age 27, moderate AR and IoT experience) in our office, equipped with 26 devices split into two groups of 13: lamps, fans, blinds, photo frames, table-height controllers and more. Conditions and device groups were counterbalanced in four orders.

  • AR condition: a Meta Quest 3 shows a white sphere next to every device; selecting it opens the generated control panel in place.
  • Non-AR condition (status quo): three smartphone apps (Philips Hue, Nanoleaf and our generated UI) and three infrared remotes for the RGB lamps, exactly one way to control each device.
  • Tasks, given by pointing: “Color the light strip red”, “Activate that fan”, “Increase the height of that table”, “Open the blinds of the left window”, “Colorize one lamp orange and the other one blue”.

User Study Results: AR Wins

  • 39 smean task time with ARvs. 68 s with apps and remotes; Wilcoxon signed-rank test, p = .003, r = −0.73
  • 84.3System Usability Scale with ARvs. 62.9 for the non-AR condition, p = .009
  • 18 / 20found AR more efficient and more fun15 found it more intuitive, 14 preferred it overall
NASA TLX box plot comparing AR and non-AR conditions.
NASA TLX (1 = very low, 21 = very high), AR in blue and non-AR in red: mental demand, effort and frustration were significantly lower with AR, while physical and temporal demand did not differ.

Participants were significantly faster with AR and reported significantly lower mental demand, effort and frustration. Qualitatively, the non-AR condition was dominated by searching for the right control, whereas AR let participants act on the object directly.

“Try every different remote and app to find which one it is for.”

P11 on the non-AR condition

“I just found the object and could change it immediately.”

P10 on the AR condition

Limitations and Next Steps

LimitationNext step
Manual steps remain: scene-graph cleanup and pre-filtering of the IoT candidate nodesNewer scene-graph pipelines already automate much of this
Mapping is type-level: identical device instances still need a human decision (guided confirmation in AR)Fuse the semantic mapping with radio-based localization or state toggling to resolve instances automatically
A single fiducial marker for spatial registration is pragmatic but does not scaleStructure-based, cross-platform anchors

Conclusion

  • Localization-ready IoT: a semantic scene graph plus LLM reasoning map network identities to physical objects, without markers or manual placement.
  • One declarative JSON specification: vendor-independent device control and automatically generated control panels for AR and smartphone clients.
  • Evaluated twice: on 100 synthetic rooms (macro F1 0.80 → 0.88 for device → object and 0.81 → 0.93 for object → device with enriched metadata) and with 20 users, who were faster, less mentally loaded and rated the AR control as more usable.

BibTeX

@article{kolberg2026loriot,
  title={Semantic Scene Graphs for Creating a Localization-Ready Internet of Things},
  author={Kolberg, Jan and Pabst, Michael and Biener, Verena and Mori, Shohei and Schmalstieg, Dieter},
  journal={IEEE Transactions on Visualization \& Computer Graphics},
  volume={32},
  number={8},
  pages={7517--7531},
  year={2026},
  doi={10.1109/TVCG.2026.3698654},
  url={https://doi.ieeecomputersociety.org/10.1109/TVCG.2026.3698654}
}