Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
The Fortieth Annual Conference on Neural Information Processing Systems , accepted (NeurIPS 2026, spotlight presentation )
Research archive · 2009–2026
I study the world in 2D and 3D, backward and forward, using deterministic and stochastic methods.
The Fortieth Annual Conference on Neural Information Processing Systems , accepted (NeurIPS 2026, spotlight presentation )
International Conference on Machine Learning , accepted (ICML 2026)
IEEE Conference on Computer Vision and Pattern Recognition , 34364--34374 (CVPR 2026)
IEEE Conference on Computer Vision and Pattern Recognition , 22028--22037 (CVPR 2026)
International Conference on Learning Representations , accepted (ICLR 2026)
IEEE Transactions on Pattern Analysis and Machine Intelligence , 2025 , accepted (TPAMI)
ACM SIGGRAPH 2025 , 122:1--122:12 (conference paper track)
IEEE Transactions on Visualization and Computer Graphics , 2025 , Vol 31 , Issue 10 , 8425--8438 (TVCG)
IEEE Transactions on Visualization and Computer Graphics , 2025 , Vol 31 , Issue 10 , 8372--8384 (TVCG)
IEEE 28th International Conference on Computer Supported Cooperative Work in Design , 316--321 (CSCWD 2025)
IEEE Transactions on Neural Networks and Learning Systems , 2025 , Vol 36 , Issue 2 , 3370--3383. (TNNLS)
The Thirty-eighth Annual Conference on Neural Information Processing Systems , 88320--88347 (NeurIPS 2024)
ACM SIGGRAPH Asia 2024 , 114:1--114:11 (conference paper track)
ACM SIGGRAPH Asia 2024 , 135:1--135:11 (conference paper track)
European Conference on Computer Vision , 18--35 (ECCV 2024)
ACM SIGGRAPH 2024 , 112:1--112:12 (conference paper track)
ACM SIGGRAPH 2024 , 113:1--113:12 (conference paper track)
ACM SIGGRAPH 2024 , 46:1--46:11 (conference paper track)
ACM SIGGRAPH 2024 , 66:1--66:9 (conference paper track)
ACM Transactions on Graphics , Vol 43 , Issue 4 , Article 83 (SIGGRAPH 2024)
Computational Visual Media , 259--278 (CVM 2024)
The 38th Annual AAAI Conference on Artificial Intelligence , 547--555 (AAAI 2024)
IEEE Transactions on Neural Networks and Learning Systems , 2024 , Vol 35 , Issue 6 , 8482--8496 (TNNLS)
IEEE Transactions on Visualization and Computer Graphics , 2024 , Vol 30 , Issue 10 , 6612--6624 (TVCG)
The Thirty-seventh Annual Conference on Neural Information Processing Systems , 31601--31620 (NeurIPS 2023)
ACM SIGGRAPH Asia 2023 , 23:1-23:11 (conference paper track)
ACM Transactions on Graphics , Vol 42 , Issue 6 , Article 244 (SIGGRAPH Asia 2023)
Computer Graphics Forum , accepted (Pacific Graphics 2023)
Computer Graphics Forum , accepted (SGP 2023)
ACM Transactions on Graphics , Vol 42 , Issue 5 , Article 169
IEEE Conference on Computer Vision and Pattern Recognition , 21726--21735 (CVPR 2023)
IEEE Conference on Computer Vision and Pattern Recognition , 12726--12735 (CVPR 2023, selected as a highlight )
IEEE Conference on Computer Vision and Pattern Recognition , 10146--10156 (CVPR 2023)
Computational Visual Media , 2023 , Vol 9 , Issue 1 , 71--85 (CVMJ)
Computational Visual Media , accepted (CVM)
Computer Graphics Forum , Vol 41 , Issue 7 , 601--610 (Pacific Graphics 2022)
Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , 2022 , Vol 6 , Issue 3 , 153:1--153:29 (IMWUT)
ACM Transactions on Graphics, Vol 41, Issue 4, Article 51 (SIGGRAPH 2022)
ACM SIGGRAPH 2022, 12:1--12:8 (conference paper track)
IEEE Conference on Computer Vision and Pattern Recognition , 20544--20554 (CVPR 2022)
IEEE Conference on Computer Vision and Pattern Recognition , 11326--11336 (CVPR 2022)
IEEE Conference on International Conference on Computer Vision , 5630--5640 (ICCV 2021)
IEEE Conference on International Conference on Computer Vision , 2753--2762 (ICCV 2021)
IEEE Conference on Computer Vision and Pattern Recognition , 13274--13283 (CVPR 2021)
ACM Transactions on Multimedia Computing, Communications, and Applications , 2021 , Vol 17 , Issue 3 , 96:1--96:17 (TOMM)
IEEE Transactions on Visualization and Computer Graphics , 2021 , Vol 27 , Issue 4 , 2298--2312 (TVCG)
With the surge of images in the information era, people demand an effective and accurate way to access meaningful visual information. Accordingly, effective and accurate communication of information has become indispensable. In this work, we propose a content-based approach that automatically generates a clear and informative visual summarization based on design principles and cognitive psychology to represent image collections. We first introduce a novel method to make representative and nonredundant summarizations of image collections, thereby ensuring data cleanliness and emphasizing important information. Then, we propose a tree-based algorithm with a two-step optimization strategy to generate the final layout that operates as follows: (1) an initial layout is created by constructing a tree randomly based on the grouping results of the input image set; (2) the layout is refined through a coarse adjustment in a greedy manner, followed by gradient back propagation drawing on the training procedure of neural networks. We demonstrate the usefulness and effectiveness of our method via extensive experimental results and user studies. Our visual summarization algorithm can precisely and efficiently capture the main content of image collections better than alternative methods or commercial tools.
The 35th Annual AAAI Conference on Artificial Intelligence , 1210--1217 (AAAI 2021)
Computational Visual Media , 2021 , Vol 7 , Issue 1 , 139--152 (CVMJ)
IEEE Transactions on Multimedia , 2020 , Vol 23 , 2794--2805 (TMM)
Art painting evaluation is sophisticated for a novice with no or limited knowledge on art criticism and history. In this study, we propose the concept of representativity to evaluate paintings instead of using professional concepts, such as genre, media, and style, which may be confusing to non-professionals. We define the concept of representativity to evaluate quantitatively the extent to which a painting can represent the characteristics of an artist’s creations. We begin by proposing a novel deep representation of art paintings, which is enhanced by style information through a weighted pooling feature fusion module. In contrast to existing feature extraction approaches, the proposed framework embeds painting styles and authorship information and learns specific artwork characteristics in a single framework. Subsequently, we propose a graph-based learning method for representativity learning, which considers intra-category and extra-category information. In view of the significance of historical factors in the art domain, we introduce the creation time of a painting into the learning process. User studies demonstrate our approach helps the public effectively access the creation characteristics of artists through sorting paintings by representativity from highest to lowest.
European Conference on Computer Vision , 90--108 (ECCV 2020)
Monocular depth estimation plays a crucial role in 3D recognition and understanding. One key limitation of existing approaches lies in their lack of structural information exploitation, which leads to inaccurate spatial layout, discontinuous surface, and ambiguous boundaries. In this paper, we tackle this problem in three aspects. First, to exploit the spatial relationship of visual features, we propose a structure-aware neural network with spatial attention blocks. These blocks guide the network attention to global structures or local details across different feature layers. Second, we introduce a global focal relative loss for uniform point pairs to enhance spatial constraint in the prediction, and explicitly increase the penalty on errors in depth-wise discontinuous regions, which helps preserve the sharpness of estimation results. Finally, based on analysis of failure cases for prior methods, we collect a new Hard Case (HC) Depth dataset of challenging scenes, such as special lighting conditions, dynamic objects, and tilted camera angles. The new dataset is leveraged by an informed learning curriculum that mixes training examples incrementally to handle diverse data distributions. Experimental results show that our method outperforms state-of-the-art approaches by a large margin in terms of both prediction accuracy on NYUDv2 dataset and generalization performance on unseen datasets.
IEEE Conference on Computer Vision and Pattern Recognition , 11207--11216 (CVPR 2020, oral presentation )
Object detection has achieved remarkable progress in the past decade. However, the detection of oriented and densely packed objects remains challenging because of following inherent reasons: (1) receptive fields of neurons are all axis-aligned and of the same shape, whereas objects are usually of diverse shapes and align along various directions; (2) detection models are typically trained with generic knowledge and may not generalize well to handle specific objects at test time; (3) the limited dataset hinders the development on this task. To resolve the first two issues, we present a dynamic refinement network which consists of two novel components, i.e., a feature selection module (FSM) and a dynamic refinement head (DRH). Our FSM enables neurons to adjust receptive fields in accordance with the shapes and orientations of target objects, whereas the DRH empowers our model to refine the prediction dynamically in an object-aware manner. To address the limited availability of related benchmarks, we collect an extensive and fully annotated dataset, namely, SKU110K-R, which is relabeled with oriented bounding boxes based on SKU110K. We perform quantitative evaluations on several publicly available benchmarks including DOTA, HRSC2016, SKU110K, and our own SKU110K-R dataset. Experimental results show that our method achieves consistent and substantial gains compared with baseline approaches. Our source code and dataset will be released to encourage follow-up research.
ACM Transactions on Graphics, Vol 39, Issue 2, Article 17 (presented at SIGGRAPH 2020)
We present a deep generative scene modeling technique for indoor environments. Our goal is to train a generative model using a feed-forward neural network that maps a prior distribution (e.g., a normal distribution) to the distribution of primary objects in indoor scenes. We introduce a 3D object arrangement representation that models the locations and orientations of objects, based on their size and shape attributes. Moreover, our scene representation is applicable for 3D objects with different multiplicities (repetition counts), selected from a database. We show a principled way to train this model by combining discriminative losses for both a 3D object arrangement representation and a 2D image-based representation. We demonstrate the effectiveness of our scene representation and the network training method on benchmark datasets. We also show the applications of this generative model in scene interpolation and scene completion.
The 34th Annual AAAI Conference on Artificial Intelligence , 5709--5716 (AAAI 2020, spotlight presentation )
Visual aesthetic assessment has been an active research field for decades. Although latest methods have achieved promising performance on benchmark datasets, they typically rely on a large number of manual annotations including both aesthetic labels and related image attributes. In this paper, we revisit the problem of image aesthetic assessment from the self-supervised feature learning perspective. Our motivation is that a suitable feature representation for image aesthetic assessment should be able to distinguish different expert-designed image manipulations, which have close relationships with negative aesthetic effects. To this end, we design two novel pretext tasks to identify the types and parameters of editing operations applied to synthetic instances. The features from our pretext tasks are then adapted for a one-layer linear classifier to evaluate the performance in terms of binary aesthetic classification. We conduct extensive quantitative experiments on three benchmark datasets and demonstrate that our approach can faithfully extract aesthetics-aware features and outperform alternative pretext schemes. Moreover, we achieve comparable results to state-of-the-art supervised methods that use 10 million labels from ImageNet.
IEEE Transactions on Multimedia , 2020 , Vol 22 , Issue 3 , 641--654 (TMM)
Real-world applications could benefit from the ability to automatically retarget an image to different aspect ratios and resolutions while preserving its visually and semantically important content. However, not all images can be equally processed. This study introduces the notion of image retargetability to describe how well a particular image can be handled by content-aware image retargeting. We propose to learn a deep convolutional neural network to rank photo retargetability, in which the relative ranking of photo retargetability is directly modeled in the loss function. Our model incorporates the joint learning of meaningful photographic attributes and image content information, which can facilitate the regularization of the complicated retargetability rating problem. To train and analyze this model, we collect a database that contains retargetability scores and meaningful image attributes assigned by six expert raters. The experiments demonstrate that our unified model can generate retargetability rankings that are highly consistent with human labels. To further validate our model, we show the applications of image retargetability in retargeting method selection, retargeting method assessment, and generating a photo collage.
International Conference on Machine Learning (ICML 2019)
In this work, we propose a novel meta-learning approach for few-shot classification, which learns transferable prior knowledge across tasks and directly produces network parameters for similar unseen tasks with training samples. Our approach, called LGM-Net, includes two key modules, namely, TargetNet and MetaNet. The TargetNet module is a neural network for solving a specific task and the MetaNet module aims at learning to generate functional weights for TargetNet by observing training samples. We also present an intertask normalization strategy for the training process to leverage common information shared across different tasks. The experimental results on Omniglot and miniImageNet datasets demonstrate that LGM-Net can effectively adapt to similar unseen tasks and achieve competitive performance, and the results on synthetic datasets show that transferable prior knowledge is learned by the MetaNet module via mapping training data to functional weights. LGM-Net enables fast learning and adaptation since no further tuning steps are required compared to other meta-learning approaches.
ACM/EG Expressive Symposium , 33--42 (Expressive 2019)
We present a learning-based style transfer algorithm for human portraits which significantly outperforms current state-of-the-art in computational overhead while still maintaining comparable visual quality. We show how to design a conditional generative adversarial network capable to reproduce the output of Fišer et al.'s patch-based method that is slow to compute but can deliver state-of-the-art visual quality. Since the resulting end-to-end network can be evaluated quickly on current consumer GPUs, our solution enables first real-time high-quality style transfer to facial videos that runs at interactive frame rates. Moreover, in cases when the original algorithmic approach of Fišer et al. fails our network can provide a more visually pleasing result thanks to generalization. We demonstrate the practical utility of our approach on a variety of different styles and target subjects.
IEEE International Conference on Multimedia and Expo , 916--921 (ICME 2019)
Taking a satisfactory picture in a low-light environment remains a challenging problem. Low-light imaging mainly suffers from noise due to the low signal-to-noise ratio. Many methods have been proposed for the task of image denoising, but they fail to work with the noise under extremely low-light conditions. Recently, deep learning based approaches have been presented that have higher objective quality than traditional methods, but they usually have high computational cost which makes them impractical to use in real-time applications or where the processing power is limited. In this paper, we propose a new residual learning based deep neural network for end-to-end extreme low-light image denoising that can not only significantly reduce the computational cost but also improve the quality over existing methods in both objective and subjective metrics. Specifically, in one setting we achieved 29x speedup with higher PSNR. Subjectively, our method provides better color reproduction and preserves more detailed texture information compared to state-of-the-art methods.
IEEE Conference on Computer Vision and Pattern Recognition , 4480--4490 (CVPR 2019, oral presentation )
We introduce a new silhouette-based representation for modeling clothed human bodies using deep generative models. Our method can reconstruct a complete and textured 3D model of a person wearing clothes from a single input picture. Inspired by the visual hull algorithm, our implicit representation uses 2D silhouettes and 3D joints of a body pose to describe the immense shape complexity and variations of clothed people. Given a segmented 2D silhouette of a person and its inferred 3D joints from the input picture, we first synthesize consistent silhouettes from novel view points around the subject. The synthesized silhouettes, which are the most consistent with the input segmentation are fed into a deep visual hull algorithm for robust 3D shape prediction. We then infer the texture of the subject's back view using the frontal image and segmentation mask as input to a conditional generative adversarial network. Our experiments demonstrate that our silhouette-based model is an effective representation and the appearance of the back view can be predicted reliably using an image-to-image translation network. While classic methods based on parametric models often fail for single-view images of subjects with challenging clothing, our approach can still produce successful results, which are comparable to those obtained from multi-view input.
IEEE Conference on Computer Vision and Pattern Recognition , 1409--1418 (CVPR 2019)
Time-lapse videos usually contain visually appealing content but are often difficult and costly to create. In this paper, we present an end-to-end solution to synthesize a time-lapse video from a single outdoor image using deep neural networks. Our key idea is to train a conditional generative adversarial network based on existing datasets of time-lapse videos and image sequences. We propose a multi-frame joint conditional generation framework to effectively learn the correlation between the illumination change of an outdoor scene and the time of the day. We further present a multi-domain training scheme for robust training of our generative models from two datasets with different distributions and missing timestamp labels. Compared to alternative time-lapse video synthesis algorithms, our method uses the timestamp as the control variable and does not require a reference video to guide the synthesis of the final output. We conduct ablation studies to validate our algorithm and compare with state-of-the-art techniques both qualitatively and quantitatively.
SIGGRAPH Asia 2018 Technical Briefs , Article 20
In this study, we present the Gourmet Photography Dataset (GPD), which is the first large-scale dataset for aesthetic assessment of food photographs. We collect 12,000 food images together with human-annotated labels (i.e., aesthetically positive or negative) to build this dataset. We evaluate the performance of several popular machine learning algorithms for aesthetic assessment of food images to verify the effectiveness and importance of our GPD dataset. Experimental results show that deep convolutional neural networks trained on GPD can achieve comparable performance with human experts in this task, even on unseen food photographs. Our experiments also provide insights to support further study and applications related to visual analysis of food images.
ACM Transactions on Graphics , Vol 37 , Issue 6 , Article 208 (SIGGRAPH Asia 2018)
Recent advances in single-view 3D hair digitization have made the creation of high-quality CG characters scalable and accessible to end-users, enabling new forms of personalized VR and gaming experiences. To handle the complexity and variety of hair structures, most cutting-edge techniques rely on the successful retrieval of a particular hair model from a comprehensive hair database. Not only are the aforementioned data-driven methods storage intensive, but they are also prone to failure for highly unconstrained input images, complicated hairstyles, and failed face detection. Instead of using a large collection of 3D hair models directly, we propose to represent the manifold of 3D hairstyles implicitly through a compact latent space of a volumetric variational autoencoder (VAE). This deep neural network is trained with volumetric orientation field representations of 3D hair models and can synthesize new hairstyles from a compressed code. To enable end-to-end 3D hair inference, we train an additional embedding network to predict the code in the VAE latent space from any input image. Strand-level hairstyles can then be generated from the predicted volumetric representation. Our fully automatic framework does not require any ad-hoc face fitting, intermediate classification and segmentation, or hairstyle database retrieval. Our hair synthesis approach is significantly more robust and can handle a much wider variation of hairstyles than state-of-the-art data-driven hair modeling techniques with challenging inputs, including photos that are low-resolution, overexposured, or contain extreme head poses. The storage requirements are minimal and a 3D hair model can be produced from an image in a second. Our evaluations also show that successful reconstructions are possible from highly stylized cartoon images, non-human subjects, and pictures taken from behind a person. Our approach is particularly well suited for continuous and plausible hair interpolation between very different hairstyles.
ACM Multimedia Conference , 879--886 (MM 2018)
Aggregation structures with explicit information, such as image attributes and scene semantics, are effective and popular for intelligent systems for assessing aesthetics of visual data. However, useful information may not be available due to the high cost of manual annotation and expert design. In this paper, we present a novel multi-patch (MP) aggregation method for image aesthetic assessment. Different from state-of-the-art methods, which augment an MP aggregation network with various visual attributes, we train the model in an end-to-end manner with aesthetic labels only (i.e., aesthetically positive or negative). We achieve the goal by resorting to an attention-based mechanism that adaptively adjusts the weight of each patch during the training process to improve learning efficiency. In addition, we propose a set of objectives with three typical attention mechanisms (i.e., average, minimum, and adaptive) and evaluate their effectiveness on the Aesthetic Visual Analysis (AVA) benchmark. Numerical results show that our approach outperforms existing methods by a large margin. We further verify the effectiveness of the proposed attention-based objectives via ablation studies and shed light on the design of aesthetic assessment systems.
European Conference on Computer Vision , 336--354 (ECCV 2018)
We present a deep learning based volumetric approach for performance capture using a passive and highly sparse multi-view capture system. State-of-the-art performance capture systems require either pre-scanned actors, large number of cameras or active sensors. In this work, we focus on the task of template-free, per-frame 3D surface reconstruction from as few as three RGB sensors, for which conventional visual hull or multi-view stereo methods fail to generate plausible results. We introduce a novel multi-view Convolutional Neural Network (CNN) that maps 2D images to a 3D volumetric field and we use this field to encode the probabilistic distribution of surface points of the captured subject. By querying the resulting field, we can instantiate the clothed human body at arbitrary resolutions. Our approach scales to different numbers of input images, which yield increased reconstruction quality when more views are used. Although only trained on synthetic data, our network can generalize to handle real footage from body performance capture. Our method is suitable for high-quality low-cost full body volumetric capture solutions, which are gaining popularity for VR and AR content creation. Experimental results demonstrate that our method is significantly more robust and accurate than existing techniques when only very sparse views are available.
International Conference on 3D Vision , 570--577 (3DV 2018)
User interaction provides useful information for solving challenging computer vision problems in practice. In this paper, we show that a very limited number of user clicks could greatly boost monocular depth estimation performance and overcome monocular ambiguities. We formulate this task as a deep structured model, in which the structured pixel-wise depth estimation has ordinal constraints introduced by user clicks. We show that the inference of the proposed model could be efficiently solved through a feed-forward network. We demonstrate the effectiveness of the proposed model on NYU Depth V2 and Stanford 2D-3D datasets. On both datasets, we achieve state-of-the-art performance when encoding user interaction into our deep models.
ACM Transactions on Graphics , Vol 36 , Issue 6 , Article 188 (SIGGRAPH Asia 2017)
3D tensor field design is important in several graphics applications such as procedural noise, solid texturing, and geometry synthesis. Different fields can lead to different visual effects. The topology of a tensor field, such as degenerate tensors, can cause artifacts in these applications. Existing 2D tensor field design systems cannot be used to handle the topology of a 3D tensor field. In this paper, we present to our knowledge the first 3D tensor field design system. At the core of our system is the ability to edit the topology of tensor fields. We demonstrate the power of our design system with applications in solid texturing and geometry synthesis.
Digital Production Symposium , Article 8 (DigiPro 2017)
In recent years the expected standard for facial animation and character performance in AAA video games has dramatically increased. The use of photogrammetric capture techniques for actor-likeness acquisition, coupled with video-based facial capture and solving methods, has improved quality across the industry. However, due to variability across project pipelines, increased per-project scope for performance capture, and a reliance on external vendors, it is often challenging to maintain visual consistency from project to project, and even from character to character within a single project. Given these factors, we identified the need for a unified, robust, and scalable pipeline for actor likeness acquisition, character art, performance capture, and character animation.
Computer Graphics Forum , Vol 36 , Issue 2 , 361--373 (Eurographics 2017, Best Paper Award Honorable Mention )
The visual richness of computer graphics applications is frequently limited by the difficulty of obtaining high-quality, detailed 3D models. This paper proposes a method for realistically transferring details (specifically, displacement maps) from existing high-quality 3D models to simple shapes that may be created with easy-to-learn modeling tools. Our key insight is to use metric learning to find a combination of geometric features that successfully predicts detail-map similarities on the source mesh, and use the learned feature combination to drive the detail transfer. The latter uses a variant of multi-resolution non-parametric texture synthesis, augmented by a high-frequency detail transfer step in texture space. We demonstrate that our technique can successfully transfer details among a variety of shapes including furniture and clothing.
IEEE Transactions on Image Processing , 2017 , Vol 26 , Issue 1 , 464--478 (TIP)
This paper presents a data-driven approach for automatically generating cartoon faces in different styles from a given portrait image. Our stylization pipeline consists of two steps: an offline analysis step to learn about how to select and compose facial components from the databases; a runtime synthesis step to generate the cartoon face by assembling parts from a database of stylized facial components. We propose an optimization framework that, for a given artistic style, simultaneously considers the desired image-cartoon relationships of the facial components and a proper adjustment of the image composition. We measure the similarity between facial components of the input image and our cartoon database via image feature matching, and introduce a probabilistic framework for modeling the relationships between cartoon facial components. We incorporate prior knowledge about image-cartoon relationships and the optimal composition of facial components extracted from a set of cartoon faces to maintain a natural, consistent and attractive look of the results. We demonstrate generality and robustness of our approach by applying it to a variety of portrait images and compare our output with stylized results created by artists via a comprehensive user study.
SIGGRAPH Asia 2016 Technical Briefs , Article 18
The design of 3D tensor fields is important in several graphics applications such as procedural noise, solid texturing, and geometry synthesis. Different fields can lead to different visual effects. The topology of a tensor field, such as degenerate tensors, can cause artifacts in these applications. Existing 2D tensor field design systems cannot handle the topology of 3D tensor fields. We present, to our best knowledge, the first 3D tensor field design system. At the core of our system is the ability to specify and control the type, number, location, shape, and connectivity of degenerate tensors. To enable such capability, we have made a number of observations of tensor field topology that were previously unreported. We demonstrate applications of our method in volumetric synthesis of solid and geometry texture as well as anisotropic Gabor noise.
SIGGRAPH Asia 2016 Technical Briefs , Article 3
We present a framework for automatically generating personalized blendshapes from actor performance measurements, while preserving the semantics of a template facial animation rig. Firstly, we capture various poses from the subject with our photogrammetry apparatus. The 3D reconstruction from each pose is then corresponded by an image-based tracking algorithm. The core of our framework is an optimization algorithm which iteratively refines the initial estimation of the blendshapes such that they can fit the performance measurements better. This framework facilitates creation of an ensemble of realistic digital-double face rigs for each individual with consistent behavior across the character set.
Similar objects are ubiquitous and abundant in both natural and artificial scenes. Determining the visual importance of several similar objects in a complex photograph is a challenge for image understanding algorithms. This study aims to define the importance of similar objects in an image and to develop a method that can select the most important instances for an input image from multiple similar objects. This task is challenging because multiple objects must be compared without adequate semantic information. This challenge is addressed by building an image database and designing an interactive system to measure object importance from human observers. This ground truth is used to define a range of features related to the visual importance of similar objects. Then, these features are used in learning-to-rank and random forest to rank similar objects in an image. Importance predictions were validated on 5922 objects. The most important objects can be identified automatically. The factors related to composition (e.g., size, location, and overlap) are particularly informative, although clarity and color contrast are also important. We demonstrate the usefulness of similar object importance on various applications, including image retargeting, image compression, image re-attentionizing, image admixture, and manipulation of blindness images.
ACM Transactions on Graphics , Vol 34 , Issue 4 , Article 125 (SIGGRAPH 2015)
Human hair presents highly convoluted structures and spans an extraordinarily wide range of hairstyles, which is essential for the digitization of compelling virtual avatars but also one of the most challenging to create. Cutting-edge hair modeling techniques typically rely on expensive capture devices and significant manual labor. We introduce a novel data-driven framework that can digitize complete and highly complex 3D hairstyles from a single-view photograph. We first construct a large database of manually crafted hair models from several online repositories. Given a reference photo of the target hairstyle and a few user strokes as guidance, we automatically search for multiple best matching examples from the database and combine them consistently into a single hairstyle to form the large-scale structure of the hair model. We then synthesize the final hair strands by jointly optimizing for the projected 2D similarity to the reference photo, the physical plausibility of each strand, as well as the local orientation coherency between neighboring strands. We demonstrate the effectiveness and robustness of our method on a variety of hairstyles and challenging images, and compare our system with state-of-the-art hair modeling algorithms.
ACM Transactions on Graphics , Vol 34 , Issue 4 , Article 47 (SIGGRAPH 2015)
There are currently no solutions for enabling direct face-to-face interaction between virtual reality (VR) users wearing head-mounted displays (HMDs). The main challenge is that the headset obstructs a significant portion of a user's face, preventing effective facial capture with traditional techniques. To advance virtual reality as a next-generation communication platform, we develop a novel HMD that enables 3D facial performance-driven animation in real-time. Our wearable system uses ultra-thin flexible electronic materials that are mounted on the foam liner of the headset to measure surface strain signals corresponding to upper face expressions. These strain signals are combined with a head-mounted RGB-D camera to enhance the tracking in the mouth region and to account for inaccurate HMD placement. To map the input signals to a 3D face model, we perform a single-instance offline training session for each person. For reusable and accurate online operation, we propose a short calibration step to readjust the Gaussian mixture distribution of the mapping before each use. The resulting animations are visually on par with cutting-edge depth sensor-driven facial performance capture systems and hence, are suitable for social interactions in virtual worlds.
IEEE Conference on Computer Vision and Pattern Recognition , 1675--1683 (CVPR 2015)
We introduce a realtime facial tracking system specifically designed for performance capture in unconstrained settings using a consumer-level RGB-D sensor. Our framework provides uninterrupted 3D facial tracking, even in the presence of extreme occlusions such as those caused by hair, hand-to-face gestures, and wearable accessories. Anyone's face can be instantly tracked and the users can be switched without an extra calibration step. During tracking, we explicitly segment face regions from any occluding parts by detecting outliers in the shape and appearance input using an exponentially smoothed and user-adaptive tracking model as prior. Our face segmentation combines depth and RGB input data and is also robust against illumination changes. To enable continuous and reliable facial feature tracking in the color channels, we synthesize plausible face textures in the occluded regions. Our tracking model is personalized on-the-fly by progressively refining the user's identity, expressions, and texture with reliable samples and temporal filtering. We demonstrate robust and high-fidelity facial tracking on a wide range of subjects with highly incomplete and largely occluded data. Our system works in everyday environments and is fully unobtrusive to the user, impacting consumer AR applications and surveillance.
ACM Transactions on Graphics , Vol 33 , Issue 6 , Article 225 (SIGGRAPH Asia 2014)
From fishtail to princess braids, these intricately woven structures define an important and popular class of hairstyle, frequently used for digital characters in computer graphics. In addition to the challenges created by the infinite range of styles, existing modeling and capture techniques are particularly constrained by the geometric and topological complexities. We propose a data-driven method to automatically reconstruct braided hairstyles from input data obtained from a single consumer RGB-D camera. Our approach covers the large variation of repetitive braid structures using a family of compact procedural braid models. From these models, we produce a database of braid patches and use a robust random sampling approach for data fitting. We then recover the input braid structures using a multi-label optimization algorithm and synthesize the intertwining hair strands of the braids. We demonstrate that a minimal capture equipment is sufficient to effectively capture a wide range of complex braids with distinct shapes and structures.
ACM Transactions on Graphics , Vol 33 , Issue 4 , Article 126 (SIGGRAPH 2014)
We introduce a data-driven hair capture framework based on example strands generated through hair simulation. Our method can robustly reconstruct faithful 3D hair models from unprocessed input point clouds with large amounts of outliers. Current state-of-the-art techniques use geometrically-inspired heuristics to derive global hair strand structures, which can yield implausible hair strands for hairstyles involving large occlusions, multiple layers, or wisps of varying lengths. We address this problem using a voting-based fitting algorithm to discover structurally plausible configurations among the locally grown hair segments from a database of simulated examples. To generate these examples, we exhaustively sample the simulation configurations within the feasible parameter space constrained by the current input hairstyle. The number of necessary simulations can be further reduced by leveraging symmetry and constrained initial conditions. The final hairstyle can then be structurally represented by a limited number of examples. To handle constrained hairstyles such as a ponytail of which realistic simulations are more difficult, we allow the user to sketch a few strokes to generate strand examples through an intuitive interface. Our approach focuses on robustness and generality. Since our method is structurally plausible by construction, we ensure an improved control during hair digitization and avoid implausible hair synthesis for a wide range of hairstyles.
Computer Graphics Forum , Vol 33 , Issue 2 , 95--104 (Eurographics 2014)
The design of video game environments, or levels, aims to control gameplay by steering the player through a sequence of designer-controlled steps, while simultaneously providing a visually engaging experience. Traditionally these levels are painstakingly designed by hand, often from pre-existing building blocks, or space templates. In this paper, we propose an algorithmic approach for automatically laying out game levels from user-specified blocks. Our method allows designers to retain control of the gameplay flow via user-specified level connectivity graphs, while relieving them from the tedious task of manually assembling the building blocks into a valid, plausible layout. Our method produces sequences of diverse layouts for the same input connectivity, allowing for repeated replay of a given level within a visually different, new environment. We support complex graph connectivities and various building block shapes, and are able to compute complex layouts in seconds. The two key components of our algorithm are the use of configuration spaces defining feasible relative positions of building blocks within a layout and a graph-decomposition based layout strategy that leverages graph connectivity to speed up convergence and avoid local minima. Together these two tools quickly steer the solution toward feasible layouts. We demonstrate our method on a variety of real-life inputs, and generate appealing layouts conforming to user specifications.
Computer Graphics Forum , Vol 33 , Issue 2 , 175--184 (Eurographics 2014)
Style transfer aims to apply the style of an exemplar model to a target one, while retaining the target's structure. The main challenge in this process is to algorithmically distinguish style from structure, a high-level, potentially ill-posed cognitive task. We recast style transfer in terms of shape analogies . In IQ testing, shape analogy queries present the subject with three shapes: source , target and exemplar , and ask them to select an output such that the transformation, or analogy , from the exemplar to the output is similar to that from the source to the target. We observe that the typical analogies we look for consist of a small set of simple transformations, which when applied to the exemplar generate a continuous, seamless output model. To assemble a shape analogy, we compute an optimal set of source-to-target transformations, such that the assembled analogy best fits these criteria. The assembled analogy is then applied to the exemplar shape to produce the desired output model. We use the proposed framework to seamlessly transfer a variety of style properties between 2D and 3D objects and demonstrate significant improvements over state of the art in style transfer. We further show that our framework can be used to successfully complete partial scans with the help of a user provided structural template, coherently propagating scan style across the completed surfaces.
ACM Transactions on Graphics , Vol 32 , Issue 4 , Article 90 (SIGGRAPH 2013)
Many natural phenomena consist of geometric elements with dynamic motions characterized by small scale repetitions over large scale structures, such as particles, herds, threads, and sheets. Due to their ubiquity, controlling the appearance and behavior of such phenomena is important for a variety of graphics applications. However, such control is often challenging; the repetitive elements are often too numerous for manual edit, while their overall structures are often too versatile for fully automatic computation. We propose a method that facilitates easy and intuitive controls at both scales: high-level structures through spatial-temporal output constraints (e.g. overall shape and motion of the output domain), and low-level details through small input exemplars (e.g. element arrangements and movements). These controls are suitable for manual specification, while the corresponding geometric and dynamic repetitions are suitable for automatic computation. Our system takes such user controls as inputs, and generates as outputs the corresponding repetitions satisfying the controls. Our method, which we call dynamic element textures , aims to produce such controllable repetitions through a combination of constrained optimization (satisfying controls) and data driven computation (synthesizing details). We use spatial-temporal samples as the core representation for dynamic geometric elements. We propose analysis algorithms for decomposing small scale repetitions from large scale themes, as well as synthesis algorithms for generating outputs satisfying user controls. Our method is general, producing a range of artistic effects that previously required disparate and specialized techniques.
ACM Transactions on Graphics , Vol 30 , Issue 4 , Article 62 (SIGGRAPH 2011)
A variety of phenomena can be characterized by repetitive small scale elements within a large scale domain. Examples include a stack of fresh produce, a plate of spaghetti, or a mosaic pattern. Although certain results can be produced via manual placement or procedural/physical simulation, these methods can be labor intensive, difficult to control, or limited to specific phenomena. We present discrete element textures, a data-driven method for synthesizing repetitive elements according to a small input exemplar and a large output domain. Our method preserves both individual element properties and their aggregate distributions. It is also general and applicable to a variety of phenomena, including different dimensionalities, different element properties and distributions, and different effects including both artistic and physically realistic ones. We represent each element by one or multiple samples whose positions encode relevant element attributes including position, size, shape, and orientation. We propose a sample-based neighborhood similarity metric and an energy optimization solver to synthesize desired outputs that observe not only input exemplars and output domains but also optional constraints such as physics, orientation fields, and boundary conditions. As a further benefit, our method can also be applied for editing existing element distributions.
Computer Graphics Forum , 2011 , Vol 30 , Issue 8 , 2156--2169 (CGF)
ACM Transactions on Graphics , Vol 28 , Issue 5 , Article 110 (SIGGRAPH Asia 2009)
A variety of animation effects such as herds and fluids contain detailed motion fields characterized by repetitive structures. Such detailed motion fields are often visually important, but tedious to specify manually or expensive to simulate computationally. Due to the repetitive nature, some of these motion fields (e.g. turbulence in fluids) could be synthesized by procedural texturing, but procedural texturing is known for its limited generality. We apply example-based texture synthesis for motion fields. Our technique is general and can take on a variety of user inputs, including captured data, manual art, and physical/procedural simulation. This data-driven approach enables artistic effects that are difficult to achieve via previous methods, such as heart shaped swirls in fluid animation. Due to the use of texture synthesis, our method is able to populate a large output field from a small input exemplar, imposing minimum user workload. Our algorithm also allows the synthesis of output motion fields not only with the same dimension as the input (e.g. 2D to 2D) but also of higher dimension, such as 3D volumetric outputs from 2D planar inputs. This cross-dimension capability supports a convenient usage scenario, i.e. the user could simply supply 2D images and our method produces a 3D motion field with similar characteristics. The motion fields produced by our method are generic, and could be combined with a variety of large-scale low-resolution motions that are easy to specify either manually or computationally but lack the repetitive structures to be characterized as textures. We apply our technique to a variety of animation phenomena, including smoke, liquid, and group motion.
Try a different title, author, venue, or year.