Skip to main navigation Skip to search Skip to main content

Spatio-Temporal Neural Rendering: Reversible Time-to-Vision Representation Learning for Human Activity Recognition

  • Zhuang Li
  • , Jing Tao
  • , Dahua Shou (Corresponding Author)

Research output: Journal article publicationJournal articleAcademic researchpeer-review

Abstract

Recently, wearable human activity recognition (WHAR) has attracted considerable attention, driven by its diverse applications in healthcare monitoring, smart environments, and human-computer interaction. Despite advances, a key challenge still limits the further development of existing WHAR methods: Individual time points contain less semantic information, and it is challenging for single-domain representation learning to reliably discern spatial-temporal (ST) dependencies hidden by inherently intricate hierarchical temporal patterns of WHAR data. To this end, we introduce a novel perspective and design innovation, Spatio-Temporal Neural Rendering (STNR), which is a vision-centric dual-domain framework tailored for WHAR that establishes reversible time-to-vision transformations between temporal data and learnable visual-like representations. Specifically, a Temporal-Visual Duality Mapping (TVDM) module encompassing a 2D rendering pathway and a 1D inverse rendering pathway is designed to unify temporal and visual modalities by leveraging their complementary strengths for enhanced WHAR. In addition, a dual-pathway Dynamic Mixing Layer (DML) is introduced to exploit both temporal variations and cross spatio-temporal patterns embedded within WHAR data. Extensive experiments conducted across six widely used WHAR datasets demonstrate that our proposed STNR achieves state-of-the-art (SOTA) performance and can be generalized to gait recognition task, providing a brand new paradigm for cross-domain time series research.

Original languageEnglish
Pages (from-to)9266-9283
JournalIEEE Transactions on Mobile Computing
Volume25
Issue number6
DOIs
Publication statusAccepted/In press - 2026

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Human activity recognition
  • representation learning
  • spatial-temporal dependency
  • time-to-vision transformations

ASJC Scopus subject areas

  • Software
  • Computer Networks and Communications
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'Spatio-Temporal Neural Rendering: Reversible Time-to-Vision Representation Learning for Human Activity Recognition'. Together they form a unique fingerprint.

Cite this