Recent highlights

  • 2026 Vahid Mahzoon and group publish a paper on learning physician-determined diagnosis from dual noisy labels, in Findings of EMNLP — extracting the clinically meaningful concept when ICD codes and physician judgment disagree. Read the paper
  • 2026 Beth Garrison and collaborators publish a paper on how AI can provide collaborative support for the unique needs of autistic job seekers, in ACM Conference on Computer-Supported Cooperative Work and Social Computing (CSCW). Read the paper
  • 2026 Saman Enayati and Elmira Talebianaraki, as shared first authors, publish a paper on automated assessment of student narratives, in Conference on Language Modeling (COLM). Read the paper
  • 2026 Sai Shi completes his Ph.D. and joins Amazon as an Applied Scientist.
  • 2026 Ashis Kumar Chanda and group publish a concordance analysis of the KDIGO definition of acute kidney injury against its coding in clinical practice, in Kidney Medicine. Read the paper
  • Ongoing NIDILRR Rehabilitation Engineering Research Center (RERC) to improve AAC technology for individuals with developmental disabilities. Temple News · Center page

Research

Research questions Three questions run through most of what the group does

Learning when good labels are scarce or unreliable

Labels may be costly, incomplete, biased, noisy, or only indirectly related to what we actually want to know. We work on weak and augmented supervision, transfer learning, few- and zero-shot learning, and learning from biased or noisy data.

Combining data with knowledge

The information needed to solve a real problem rarely lives in one clean source. It may be spread across text, structured records, measurements, knowledge bases, simulations, graphs, trajectories, and expert knowledge. We study multimodal, representation, and knowledge-aware learning.

Complementary human and machine intelligence

People contribute expertise, context, judgment, goals, and lived experience that a predictive model alone cannot supply. We study human-in-the-loop and collaborative AI systems in which machine learning augments human capability rather than replacing human participation.

Representative recent projects Where those questions are being pursued now

Human capability & accessibility

How can AI expand communication, employment, and participation while preserving human agency?

Neurodiversity, augmentative and alternative communication, employment support, assistive technology, personalization, and human–AI collaboration — interdisciplinary research with HCI, behavioral science, special education, communication sciences, and disability research experts. The goal is not to automate a task but to understand how AI can support people within the settings where the technology is actually used.

Health & biomedical discovery

How can AI extract meaningful knowledge from biomedical information that is heterogeneous and incomplete?

Electronic health records, patient-provider communication, biomedical information extraction, population health, genomics, and biological sequence analysis. A recurring challenge is that the concept we care about is often not directly observable: diagnoses may be inconsistently coded, evidence distributed across text and structured records, and expert interpretation essential.

Science, engineering & intelligent systems

How can models learn when observations are expensive, limited, or collected under changing conditions?

Scientific machine learning, wireless systems, cybersecurity, robotics, materials, simulation, sensing, and remote sensing. We are particularly interested in data efficiency: how existing knowledge, simulation, transfer learning, multiple sources of information, and adaptive learning can reduce the need for costly new observations.

Contributions

Computational biology and bioinformatics

For more than two decades I have developed machine-learning methods that infer protein structure, function, and phenotype from complex and incomplete molecular data: protein intrinsic disorder, protein function prediction, integration of heterogeneous biological data, generative models of protein sequences, and genomic determinants of biological traits. This work is done in close collaboration with experimental and computational biologists.

Benchmarking and rigorous evaluation

A model is only useful if its performance is robust, reproducible, and holds on independent data. I have taken part in community-wide blind assessments, of protein disorder prediction at CASP and of protein function prediction at CAFA, and designed evaluation frameworks in other domains, studying training and testing methodology, ablation, sampling bias, and human evaluation of language-model outputs.

Learning from heterogeneous, limited, and imperfect data

Scientific problems rarely come with homogeneous, independently sampled, fully labeled data. I have developed methods for multi-source, multi-task, and semi-supervised learning, representation and structured learning, learning from biased data, and scalable learning under memory and compute budgets, applied in bioinformatics, medicine, remote sensing, and engineering.

NLP and AI for biomedical information

Much of what matters in biomedicine is written in text. I have worked on representation learning for medical concepts, prediction from electronic health record narratives, biomedical relation extraction, medical code prediction with external knowledge, and radiology report simplification, with growing emphasis on understanding what language models can and cannot do through systematic and human evaluation.

Human-centered and human-in-the-loop AI

Where human judgment, expertise, and values remain essential, AI should complement people rather than replace them. My recent work builds and evaluates such systems in the settings where they are used, especially assistive technology for people with developmental disabilities, including AI-assisted augmentative and alternative communication and employment support, together with researchers in disability, communication, psychology, education, and HCI.

How we work

Most projects begin with a difficult real-world problem rather than an existing benchmark.

We work with domain experts to understand the problems they care about, how the data were created, what is known and what information is missing, and why existing approaches fall short. Finding answers frequently requires new machine-learning approaches and algorithms.

Understand the problem → identify the learning challenge → develop new methods → build and evaluate a system → produce methodological and domain insight

Applications are where the most interesting machine-learning problems originate. The way we work has led us to collaborations across medicine, biology, public health, communication sciences, behavioral science, education, materials science, remote sensing, robotics, and industrial engineering.

Center for Hybrid Intelligence. Beyond my research group, I direct Temple's CHI, which extends this model of collaborative research beyond a single laboratory — connecting AI researchers, domain experts, students, and external partners. Visit CHI →

Group

Current doctoral students

Zhuoan Zhou

Human-computer interaction, AI-assisted systems.

Vahid Mahzoon

Large language models, applied machine learning, medical informatics.

Elmira Talebianaraki

Human-computer interaction, medical informatics, natural language processing.

Students are encouraged to develop methodological depth while engaging seriously with substantive application areas. Some projects are primarily methodological; others are deeply interdisciplinary; many combine both. See Work with us if you are considering doctoral study.

Doctoral alumni Five most recent shown; expand for the full list

NamePh.D.PositionLinks
Sai Shi2026Applied Scientist, AmazonScholar
Elizabeth Garrison2025Startup founderScholar
Hanzi Xu2024Research Scientist, NetflixScholar
Ziyu Yang2024Software Engineer, GoogleScholar
Sandro Hauri2022Machine Learning Engineer, Swiss Federal RailwaysScholar
Show all 16 doctoral alumni
NamePh.D.PositionLinks
Ashis Chanda2022Senior Data Scientist, American ExpressScholar
Aniruddha Maiti2022Assistant Professor, West Virginia State UniversityScholar
Maxim Shapovalov2019Postdoctoral Researcher, Fox Chase Cancer Center
Tian Bai2019Software Engineer, GoogleScholar
Shanshan Zhang2019Senior Research Scientist, MetaScholar
Vladimir Coric2014Lead Data Scientist, SEI Investments
Nemanja Djuric2013Staff Engineer / Tech Lead Manager, AuroraScholar
Liang Lan2012Professor, Hong Kong Baptist UniversityScholar
Mihajlo Grbovic2012Principal Machine Learning Scientist, AirbnbScholar
Vuk Malbasa2011Assistant Professor, University of Novi SadScholar
Zhuang John Wang2010Staff Software Engineer, Meta

Representative publications

Click any entry to read its abstract. Complete record below — Google Scholar · DBLP

Journal 2026 Concordance Analysis Between the KDIGO Definition of Acute Kidney Injury and Its Coding in Clinical PracticeKidney Medicine, 2026

Rationale & Objective Acute kidney injury (AKI) in research is typically identified using KDIGO criteria based on changes in serum creatinine (SCr) levels and urine output volumes, whereas in clinical practice, AKI is documented through International Classification of Diseases, Ninth or Tenth Revision (ICD-9/10) diagnosis codes at discharge. This study evaluated how KDIGO criteria-defined AKI corresponds to clinically documented AKI and analyzed discordant cases to understand the respective strengths of each method. Study Design Retrospective cohort study. Setting & Participants We analyzed 54,210 hospital admissions with intensive care unit stays from 43,530 patients at Beth Israel Deaconess Medical Center over 15 years (2008-2022). Predictors SCr levels and urine output volumes recorded during a hospital admission. Outcomes Hospital-coded AKI, defined as the presence of AKI among the ICD-9/10 codes assigned at discharge. Analytical Approach We quantified concordance between KDIGO-defined AKI and hospital-coded AKI using Cohen’s kappa (κ) and examined patterns of disagreement. Machine learning models were trained to assess the association between SCr level and hospital-coded AKI. Results KDIGO criteria-defined AKI had low concordance with ICD-coded AKI (κ = 0.37), whereas a formula using the maximum SCr level achieved higher concordance (κ = 0.61). KDIGO criteria frequently detected small increases in SCr levels that were not coded, whereas coding frequently detected patients with high SCr levels at presentation that were missed using KDIGO criteria. Limitations Our analysis is confined to a single hospital system in the United States. Data used in this study lacked information about preadmission SCr baseline levels. Conclusions Further work is needed to clarify how KDIGO criteria are used in clinical practice and to refine how AKI is identified.

Conference 2026 Learning to Extract Physician-Determined Diagnosis from Clinical Narratives Using Dual Noisy LabelsFindings of EMNLP, 2026

Electronic health records (EHRs) offer valuable opportunities for retrospective studies across a wide range of diseases. Many retrospective studies require reliable diagnosis labels for each patient but EHR datasets often do not provide precise labels that reflect the physician-determined diagnosis. Therefore, existing NLP approaches in retrospective studies typically rely on diagnosis proxies such as ICD codes, which are often noisy and may not reflect the physician-determined diagnosis. In this work, we propose a dual noisy label learning approach based on a latent variable framework to train a deep learning model that extracts physician-determined diagnoses from discharge summaries by leveraging two complementary noisy supervision signals: ICD diagnosis codes and discharge diagnoses. We evaluate our methodology on Acute Kidney Injury (AKI) diagnosis as a challenging case using a gold-standard dataset created through expert adjudication. Our approach substantially outperforms commonly used proxies and our interpretability analysis shows that the model’s attention aligns closely with expert clinical reasoning. This approach can be applied for extraction of arbitrary physician-determined diagnoses and allows generating reliable diagnosis labels from clinical narratives without requiring manual annotation.

Conference 2026 Enhancing Interview Preparation for Autistic Job Seekers through Human-AI CollaborationACM CSCW, 2026

Autistic adults commonly face challenges with social interaction, making job interviews a particularly difficult part of the job search process. This work seeks to understand how AI can provide collaborative support to address the unique needs of autistic job seekers. First, we conducted a formative study with five autistic job seekers and five vocational coaches to gain an in-depth understanding of specific challenges during interview preparation for autistic job seekers, and also discuss their perspectives on the design and use of an intelligent interview coach. In this formative study, we identified the unique challenges faced by autistic job seekers (masking and disclosure) and learned the importance of concise feedback delivery and procedural scaffolding when preparing autistic job seekers for interviews. In our second study, we developed a design probe for an intelligent interview coaching system tailored to autistic job seekers and conducted this follow-up study with three autistic job seekers and three vocational coaches to explore the design concepts identified in the formative study. In this study, we learned that adding anxiety and confidence assessments, personalization components, and visual structure should be design considerations when developing an intelligent interview coaching system for autistic job seekers. From the results of this study, we also discuss considerations for human-AI collaboration within intelligent interview coaching systems. We highlight the potential and limitations of intelligent interview coaching systems in supporting autistic job seekers and vocational coaches and offer design considerations for developing future collaborative AI assistants for interview preparation.

Conference 2025 Deep Learning-Based Pedestrian Simulation with Limited Real-World Training DataIJCAI, 2025

Simulating pedestrian movement is important for applications such as disaster management, robotics, and game design. While deep learning models have been extensively used on related problems, their use as pedestrian simulators remains relatively unexplored. This paper aims to encourage more research in this direction in two ways. First, it proposes an evaluation framework that is applicable to both traditional and deep learning based simulators. Second, it proposes and evaluates several ideas related to input representation, choice of neural architecture, exploiting knowledge-based simulators in data poor regimes, and repurposing trajectory prediction models. Our extensive experiments provide several useful insights for future research in pedestrian simulation. The code is available at https://github.com/vmahzoon76/DL-Crowd-Sim.

Journal 2025 Evolutionary Sparse Learning Reveals the Shared Genetic Basis of Convergent TraitsNature Communications, 2025

Cases abound in which nearly identical traits have appeared in distant species facing similar environments. These unmistakable examples of adaptive evolution offer opportunities to gain insight into their genetic origins and mechanisms through comparative analyses. Here, we present an approach to build genetic models that underlie the independent origins of convergent traits using evolutionary sparse learning with paired species contrast (ESL-PSC). We tested the hypothesis that common genes and sites are involved in the convergent evolution of two key traits: C4 photosynthesis in grasses and echolocation in mammals. Genetic models were highly predictive of independent cases of convergent evolution of C4 photosynthesis. Genes contributing to genetic models for echolocation were highly enriched for functional categories related to hearing, sound perception, and deafness, a pattern that has eluded previous efforts applying standard molecular evolutionary approaches. These results support the involvement of sequence substitutions at common genetic loci in the evolution of convergent traits. Benchmarking on empirical and simulated datasets showed that ESL-PSC could be more sensitive in proteome-scale analyses to detect genes with convergent molecular evolution associated with the acquisition of convergent traits. We conclude that phylogeny-informed machine learning naturally excludes apparent molecular convergences due to shared species history, enhances the signal-to-noise ratio for detecting molecular convergence, and empowers the discovery of common genetic bases of trait convergences.

Conference 2024 X-Shot: Frequent-, Few-, and Zero-Shot Classification in One SystemFindings of ACL, 2024

In recent years, few-shot and zero-shot learning, which learn to predict labels with limited annotated instances, have garnered significant attention. Traditional approaches often treat frequent-shot (freq-shot; labels with abundant instances), few-shot, and zero-shot learning as distinct challenges, optimizing systems for just one of these scenarios. Yet, in real-world settings, label occurrences vary greatly. Some of them might appear thousands of times, while others might only appear sporadically or not at all. For practical deployment, it is crucial that a system can adapt to any label occurrence. We introduce a novel classification challenge: X-shot, reflecting a real-world context where freq-shot, few-shot, and zero-shot labels co-occur without predefined limits. Here, X can span from 0 to positive infinity. The crux of X-shot centers on open-domain generalization and devising a system versatile enough to manage various label scenarios. To solve X-shot, we propose BinBin (Binary INference Based on INstruction following) that leverages the Indirect Supervision from a large collection of NLP tasks via instruction following, bolstered by Weak Supervision provided by large language models. BinBin surpasses previous state-of-the-art techniques on three benchmark datasets across multiple domains. To our knowledge, this is the first work addressing X-shot learning, where X remains variable.

Conference 2024 Two-Pronged Human Evaluation of ChatGPT Self-Correction in Radiology Report SimplificationFindings of ACL, 2024

Radiology reports are highly technical documents aimed primarily at doctor-doctor communication. There has been an increasing interest in sharing those reports with patients, necessitating providing them patient-friendly simplifications of the original reports. This study explores the suitability of large language models in automatically generating those simplifications. We examine the usefulness of chain-of-thought and self-correction prompting mechanisms in this domain. We also propose a new evaluation protocol that employs radiologists and laypeople, where radiologists verify the factual correctness of simplifications, and laypeople assess simplicity and comprehension. Our experimental results demonstrate the effectiveness of self-correction prompting in producing high-quality simplifications. Our findings illuminate the preferences of radiologists and laypeople regarding text simplification, informing future research on this topic.

Journal 2024 Leveraging Communication Partner Speech to Automate Augmented Input for Minimally Verbal Autistic ChildrenAmerican Journal of Speech-Language Pathology, 2024

Purpose: Augmentative and alternative communication (AAC) technology innovation is urgently needed to improve outcomes for children on the autism spectrum who are minimally verbal. One potential technology innovation is applying artificial intelligence (AI) to automate strategies such as augmented input to increase language learning opportunities while mitigating communication partner time and learning barriers. Innovation in AAC research and design methodology is also needed to empirically explore this and other applications of AI to AAC. The purpose of this report was to describe (a) the development of an AAC prototype using a design methodology new to AAC research and (b) a preliminary investigation of the efficacy of this potential new AAC capability. Method: The prototype was developed using a Wizard-of-Oz prototyping approach that allows for initial exploration of a new technology capability without the time and effort required for full-scale development. The preliminary investigation with three children on the autism spectrum who were minimally verbal used an adapted alternating treatment design to compare the effects of a Wizard-of-Oz prototype that provided automated augmented input (i.e., pairing color photos with speech) to a standard topic display (i.e., a grid display with line drawings) on visual attention, linguistic participation, and (for one participant) word learning during a circle activity. Results: Preliminary investigation results were variable, but overall participants increased visual attention and linguistic participation when using the prototype. Conclusions: Wizard-of-Oz prototyping could be a valuable approach to spur much needed innovation in AAC. Further research into efficacy, reliability, validity, and attitudes is required to more comprehensively evaluate the use of AI to automate augmented input in AAC.

Conference 2023 Group Activity Recognition in Basketball Tracking Data (NETS)ECAI, 2023

Like many team sports, basketball involves two groups of players who engage in collaborative and adversarial activities to win a game. Players and teams are executing various complex strategies to gain an advantage over their opponents. Defining, identifying, and analyzing different types of activities is an important task in sports analytics, as it can lead to better strategies and decisions by the players and coaching staff. The objective of this paper is to automatically recognize basketball group activities from tracking data representing locations of players and the ball during a game. We propose a novel deep learning approach for group activity recognition (GAR) in team sports called NETS. To efficiently model the player relations in team sports, we combined a Transformer-based architecture with LSTM embedding, and a team-wise pooling layer to recognize the group activity. Training such a neural network generally requires a large amount of annotated data, which incurs high labeling cost. To alleviate this problem, we pretrain the neural network on a self-supervised trajectory prediction task and fine-tune it using a mix of strong and weak labels. We used a large tracking data set from 632 NBA games to evaluate our approach. The results show that NETS is capable of learning group activities with high accuracy, and that self- and weak-supervised training in NETS have a positive impact on GAR accuracy.

Journal 2021 Structure Motif-Centric Learning Framework for Inorganic Crystalline SystemsScience Advances, 2021

Incorporation of physical principles in a machine learning (ML) architecture is a fundamental step toward the continued development of artificial intelligence for inorganic materials. As inspired by the Pauling’s rule, we propose that structure motifs in inorganic crystals can serve as a central input to a machine learning framework. We demonstrated that the presence of structure motifs and their connections in a large set of crystalline compounds can be converted into unique vector representations using an unsupervised learning algorithm. To demonstrate the use of structure motif information, a motif-centric learning framework is created by combining motif information with the atom-based graph neural networks to form an atom-motif dual graph network (AMDNet), which is more accurate in predicting the electronic structures of metal oxides such as bandgaps. The work illustrates the route toward fundamental design of graph neural network learning architecture for complex materials by incorporating beyond-atom physical principles.

Journal 2021 Generative Capacity of Probabilistic Protein Sequence ModelsNature Communications, 2021

Potts models and variational autoencoders (VAEs) have recently gained popularity as generative protein sequence models (GPSMs) to explore fitness landscapes and predict the effect of mutations. Despite encouraging results, quantitative characterization and comparison of GPSM-generated probability distributions is still lacking. It is currently unclear whether GPSMs can faithfully reproduce the complex multi-residue mutation patterns observed in natural sequences arising due to epistasis. We develop a set of sequence statistics to comparatively assess the accuracy, or “generative capacity”, of three GPSMs: a pairwise Potts Hamiltonian, a vanilla VAE, and a site-independent model, using natural and synthetic datasets. We show that the generative capacity of the Potts Hamiltonian model is the largest; the higher order mutational statistics generated by the model agree with those observed for natural sequences. In contrast, we show that the vanilla VAE’s generative capacity lies between the pairwise Potts and site-independent models. Importantly, our work measures GPSM generative capacity in terms of higher-order sequence covariation and provides a new framework for evaluating and interpreting GPSM accuracy that emphasizes the role of epistasis.

Show all publications, 2000–2026

Complete verified record of journal and conference publications, newest first. Maintained independently of Google Scholar so the list does not depend on an external service. Each entry expands to its abstract and links to the publisher.

2026

Conference Mahzoon, V., Talebianaraki, E., Ojeniyi, S., Efferen, T., Koul, S., Vucetic, Z., Gillespie, A., Vucetic, S., Learning to Extract Physician-Determined Diagnosis from Clinical Narratives Using Dual Noisy Labels, Findings of the Association for Computational Linguistics: EMNLP, 2026.

Electronic health records (EHRs) offer valuable opportunities for retrospective studies across a wide range of diseases. Many retrospective studies require reliable diagnosis labels for each patient but EHR datasets often do not provide precise labels that reflect the physician-determined diagnosis. Therefore, existing NLP approaches in retrospective studies typically rely on diagnosis proxies such as ICD codes, which are often noisy and may not reflect the physician-determined diagnosis. In this work, we propose a dual noisy label learning approach based on a latent variable framework to train a deep learning model that extracts physician-determined diagnoses from discharge summaries by leveraging two complementary noisy supervision signals: ICD diagnosis codes and discharge diagnoses. We evaluate our methodology on Acute Kidney Injury (AKI) diagnosis as a challenging case using a gold-standard dataset created through expert adjudication. Our approach substantially outperforms commonly used proxies and our interpretability analysis shows that the model’s attention aligns closely with expert clinical reasoning. This approach can be applied for extraction of arbitrary physician-determined diagnoses and allows generating reliable diagnosis labels from clinical narratives without requiring manual annotation.

Conference Enayati, S., Talebianaraki, E., Yang, Z., Vucetic, S., Spencer, T.D., Benchmarking of Automated Assessment of K-5 Student Narratives Using Large Language Models, Proceedings of the 3rd Annual Conference on Language Modeling (COLM), 2026.

Assessing narrative language is essential for evaluating early literacy development, but manual scoring of children’s narratives is labor-intensive and requires expert training. In this work, we study whether large language models (LLMs) can reliably assess narrative structure in K–5 student stories, focusing on rubric-based scoring of discourse elements such as character, setting, and problem. We present a systematic evaluation of LLM-based scoring across two expert-annotated datasets of transcribed oral narratives, comparing fine-tuning and few-shot prompting approaches across a range of modern LLM architectures. Using human inter-rater reliability (IRR) as a reference, we find that even smaller-sized fine-tuned LLMs achieve near-human Quadratic Weighted Kappa (QWK) accuracy on in-distribution narratives, while few-shot prompting is less accurate. However, fine-tuned models do not generalize robustly and exhibit substantial degradation on out-of-sample story types. We find that few-shot prompting is competitive in low-resource settings, while fine-tuning is preferable when sufficient labeled data is available. This underscores a need for further research in low-resource automated narrative assessment.

Conference Garrison, E., MacNeil, S., Hong, S.R., Ara, Z., Hantula, D., Vucetic, S., Enhancing Interview Preparation for Autistic Job Seekers through Human-AI Collaboration, Proceedings of the 29th ACM Conference on Computer-Supported Cooperative Work and Social Computing (CSCW), 2026.

Autistic adults commonly face challenges with social interaction, making job interviews a particularly difficult part of the job search process. This work seeks to understand how AI can provide collaborative support to address the unique needs of autistic job seekers. First, we conducted a formative study with five autistic job seekers and five vocational coaches to gain an in-depth understanding of specific challenges during interview preparation for autistic job seekers, and also discuss their perspectives on the design and use of an intelligent interview coach. In this formative study, we identified the unique challenges faced by autistic job seekers (masking and disclosure) and learned the importance of concise feedback delivery and procedural scaffolding when preparing autistic job seekers for interviews. In our second study, we developed a design probe for an intelligent interview coaching system tailored to autistic job seekers and conducted this follow-up study with three autistic job seekers and three vocational coaches to explore the design concepts identified in the formative study. In this study, we learned that adding anxiety and confidence assessments, personalization components, and visual structure should be design considerations when developing an intelligent interview coaching system for autistic job seekers. From the results of this study, we also discuss considerations for human-AI collaboration within intelligent interview coaching systems. We highlight the potential and limitations of intelligent interview coaching systems in supporting autistic job seekers and vocational coaches and offer design considerations for developing future collaborative AI assistants for interview preparation.

Conference Feng, H., Ara, Z., Hundt, A., Vucetic, S., Joon Young Chung, J., Hong, S.R., Lost in Translation: Understanding Autistic–Neurotypical Communication Style Differences in Job Postings. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI) (pp. 1-17), 2026.

Autistic adults often use different communication styles than neurotypical individuals (NTs). While prior research has documented how such gaps disadvantage autistic job seekers, no study has systematically examined when these differences arise in language use and why autistic adults encounter interpretive gaps. This work seeks to datafy and characterize these communication challenges. We built an annotation interface and recruited 20 autistic adults to analyze 10 job postings each that they had selected as cases where they felt “lost in translation.” Participants annotated text spans using six categories informed by speech and language literature: unclear, ambiguous, incomplete, inappropriate, negative, and other. Follow-up interviews showed that lexical difficulties were rarely barriers; rather, challenges stemmed from interpreting implicit social arrangements or unstated expectations. We release the anonymized annotation data as the first-of-its-kind dataset documenting autistic–NT communication style differences. We conclude with implications for designing supports that foster clearer autistic–NT communication.

Journal Chanda, A.K., Talebianaraki, E., Mahzoon, V., Vucetic, Z., Vucetic, S., Concordance Analysis Between KDIGO Definition of Acute Kidney Injury and Its Coding in Clinical Practice. Kidney Medicine, p.101374, 2026.

Rationale & Objective Acute kidney injury (AKI) in research is typically identified using KDIGO criteria based on changes in serum creatinine (SCr) levels and urine output volumes, whereas in clinical practice, AKI is documented through International Classification of Diseases, Ninth or Tenth Revision (ICD-9/10) diagnosis codes at discharge. This study evaluated how KDIGO criteria-defined AKI corresponds to clinically documented AKI and analyzed discordant cases to understand the respective strengths of each method. Study Design Retrospective cohort study. Setting & Participants We analyzed 54,210 hospital admissions with intensive care unit stays from 43,530 patients at Beth Israel Deaconess Medical Center over 15 years (2008-2022). Predictors SCr levels and urine output volumes recorded during a hospital admission. Outcomes Hospital-coded AKI, defined as the presence of AKI among the ICD-9/10 codes assigned at discharge. Analytical Approach We quantified concordance between KDIGO-defined AKI and hospital-coded AKI using Cohen’s kappa (κ) and examined patterns of disagreement. Machine learning models were trained to assess the association between SCr level and hospital-coded AKI. Results KDIGO criteria-defined AKI had low concordance with ICD-coded AKI (κ = 0.37), whereas a formula using the maximum SCr level achieved higher concordance (κ = 0.61). KDIGO criteria frequently detected small increases in SCr levels that were not coded, whereas coding frequently detected patients with high SCr levels at presentation that were missed using KDIGO criteria. Limitations Our analysis is confined to a single hospital system in the United States. Data used in this study lacked information about preadmission SCr baseline levels. Conclusions Further work is needed to clarify how KDIGO criteria are used in clinical practice and to refine how AKI is identified.

2025

Conference Mahzoon, V., Liu, A., Vucetic, S., Deep Learning-Based Pedestrian Simulation with Limited Real-World Training Data: An Evaluation Framework. Proceedings of the 34th International Joint Conference on Artificial Intelligence (IJCAI), Montreal, Canada, 2025.

Simulating pedestrian movement is important for applications such as disaster management, robotics, and game design. While deep learning models have been extensively used on related problems, their use as pedestrian simulators remains relatively unexplored. This paper aims to encourage more research in this direction in two ways. First, it proposes an evaluation framework that is applicable to both traditional and deep learning based simulators. Second, it proposes and evaluates several ideas related to input representation, choice of neural architecture, exploiting knowledge-based simulators in data poor regimes, and repurposing trajectory prediction models. Our extensive experiments provide several useful insights for future research in pedestrian simulation. The code is available at https://github.com/vmahzoon76/DL-Crowd-Sim.

Conference Shi, S., Mahzoon, V., Wang, X., Mao, S., Wu, J., Vucetic, S., Towards a Unified Few-Shot Learning Evaluation Framework for RF Fingerprinting. Proceedings of the 34th International Conference on Computer Communications and Networks (ICCCN), Tokyo, Japan, 2025.

Radio frequency (RF) fingerprinting is a technique used to identify a wireless device based on its specific and unique hardware characteristics. In recent years, deep learning has been utilized for RF fingerprinting due to its superiority in feature extraction and higher classification accuracy. However, one major challenge of deep learning-based RF fingerprinting is that wireless signals are highly sensitive to environmental conditions, causing the device fingerprints captured in one environment to not transfer well to another. Hence, deep learning models are found to perform well in the same condition but lose their ability to classify devices in the new condition. In this paper, we examine three transfer learning techniques to mitigate the domain shift problem in RF fingerprinting and compare them with two well-defined baselines. The three RF fingerprinting datasets under various scenarios are examined to explore how environmental factors impact RF fingerprinting, such as transmitter locations, transmitter distance, and device configurations. We identify the most challenging scenarios and study how environmental factors lead to model deterioration through t-SNE visualization.

Conference Niknami, N., Mahzoon, V., Biswas, R., Vucetic, S., Wu, J., An Interpretable Multi-Modal Transformer-Based Intrusion Detection System Utilizing Log and PCAP, Proceedings of the 22nd IEEE International Conference on Mobile Ad-Hoc and Smart Systems (MASS), 2025.

Intrusion detection systems (IDS) primarily rely on signature-based approaches, which can fail to detect novel or sophisticated attacks. This paper addresses the underutilized potential of leveraging a multi-modal approach that combines packet capture (PCAP) and log data for anomaly detection. To enhance detection capabilities, we propose an interpretable hybrid neural network architecture, TransIDS, that integrates a packet-based transformer with an efficient transformer-based language model for log messages. The proposed framework extracts semantic vectors from raw log messages and concatenates them with packet embeddings. An attention-based classification model then detects anomalies by determining the importance of each log message and packet for the neural network’s decision. By fusing spatial features from PCAP data with temporal features from log data, TransIDS utilizes this multi-modal data fusion to identify anomalies that might be missed by conventional systems. This approach not only leverages the strengths of two distinct transformer-based architectures but also provides a more comprehensive analysis of network traffic, leading to more effective detection of previously undetected attacks and strengthening overall network security. We use a real testbed for our experiments to validate the effectiveness of our proposed approach.

Conference Garrison, E., MacNeil, S., Lorah, E.R., Holyfield, C., Vucetic, S., Exploring Engagement Opportunities for Autistic Children: Using AAC as a Controller in a Wizard-of-Oz Coloring Game. Proceedings of the ACM on Human-Computer Interaction (HCI), 9(1), pp.1-21, 2025.

Autistic children face significant challenges in vocal communication and social interaction, often leading to social isolation. There is evidence that Augmentative and Alternative Communication (AAC) offers support to mitigate these challenges, enabling them to communicate with non-vocal means through forms of AAC, such as speech-generation devices (SGDs). However, the adoption and use of SGDs are hindered by several factors, including the large amount of practice required to learn to use SGDs and the limited options for highly engaging social learning contexts. Our study introduces the novel approach of using SGDs as game controller for digital and interactive games. With three design goals guiding our work, we conducted a Wizard-of-Oz formative case study with five participants aged 3-5 years, who were learning to use their SGD. We simulated a digital coloring game, integrating the speech-generated output of the participant's SGD to function as the game's controller. From this case study, we observed that all participants engaged with the game using their SGD for at least one turn, and two participants also engaged in emerging joint attention responses with the game and game's facilitator. This paper discusses these findings and contributes directions for future research, with suggestions for the design of future SGD-controlled games and exploration of social connection and collaboration between autistic children who use AAC and their caregivers, siblings, and peers.

Journal Garrison, E., MacNeil, S., Hantula, D.A., West, M., Dragut, E., Tincani, M., Vucetic, S., Exploring the challenges and assistive technology for autistic job seekers across employment pathways. Research in Developmental Disabilities, 167, p.105155, 2025.

Background: When transitioning from high school, autistic job seekers often navigate three different pathways to employment: University, Job Coaching, and Self-Directed (defined as those job seekers who independently complete the job search process, without formal support). Assistive technology may aid job seekers throughout the job seeking process. The aim of this study is to learn more about the challenges and assistive technology that autistic job seekers encounter while navigating these three different employment pathways. Methods: Qualitative semi-structured interviews were conducted with fifteen stakeholders in the United States, autistic job seekers and support personnel, within each pathway of the hiring process to gather information regarding the challenges autistic job seekers encounter, and the assistive technology they use to address those challenges. Results: From a thematic analysis of these interviews, we found that autistic job seekers along each pathway commonly move through the following phases of the hiring process or “checkpoints”: resume building, networking, job search, job application, and interviews. Autistic job seekers also face challenges within each checkpoint, such as knowing when and what to disclose; self-efficacy, anxiety, and communication challenges; and a lack of communication from potential employers. We also learned that some self-directed autistic job seekers, when compared to those in the University and Job Coaching pathways, may not be using assistive technologies available in the job search process. From our interviews, we also learned the types of assistive technology that autistic job seekers and assistants use in the job seeking process which can be classified as organizational tools, connectivity tools, and visual media tools. Conclusion and implications: Our findings revealed a necessity to connect self-directed autistic job seekers to assistive technology available. Based on these results, we present suggestions for future research and design suggestions for developing assistive technology for autistic job seekers.

Journal West, M., Tincani, M., Hantula, D., Hong, S.R., Vucetic, S., Dragut, E., Applying data science practices to identify characteristics of postsecondary autism support programs from their websites. Journal of Autism and Developmental Disorders, pp.1-47, 2025.

Purpose: An increasing number of autistic students in the United States are seeking post-secondary education. In response, some post-secondary institutions have established Autism Support Programs (ASP) to address the comprehensive needs of this population. There is little up-to-date, comprehensive information about which institutions host these programs, what types of services they offer, and what is required to access them. Methods: Expanding on previous research, we introduce a new method, which utilizes established data science techniques, to identify ASPs at post-secondary institutions in the U.S. Our technique also allows us to identify the characteristics of the ASPs, including admissions requirements, cost, structure, and supports offered. Results: Results highlight our method is more efficient and more robust than previous methods from the literature. For example, we identify 49 schools hosting ASPs that were not identified in past literature searches. We report on the characteristics of identified ASPs such as application process, most common supports and program cost. Conclusion: The bi-directional change in the number of ASPs shows that this is an evolving field, requiring automated tools to enable regular updates to data. Although it is promising that a relative handful of U.S. schools have established these programs, a large majority of post-secondary institutions have not, and for those that host them, barriers to access exist, including the necessity of an ASD diagnosis, coupled with up-front and ongoing costs.

Journal Holyfield, C., O’Neill Zimmerman, T., MacNeil, S., Sammarco Caldwell, N., Patel, P., Griffen, B., Lorah, E., Dragut, E., Vucetic, S., Preliminary Investigation of Context-Aware Augmentative and Alternative Communication with Automated Just-in-Time Cloze Phrase Response Options for Social Participation from Children on the Autism Spectrum. Folia Phoniatrica et Logopaedica, 77(3), pp.269-283, 2025.

Introduction: Social participation for emerging symbolic communicators on the autism spectrum is often restricted. This is due in part to the time and effort required for both children and partners to use traditional augmentative and alternative communication (AAC) technologies during fast-paced social routines. Innovations in artificial intelligence provide the potential for context-aware AAC technology that can provide just-in-time communication options based on linguistic input from partners to minimize the time and effort needed to use AAC technologies for social participation. Methods: This preliminary study used an alternating treatment design to compare the effects of a context-aware AAC prototype with automated cloze phrase response options to traditional AAC for supporting three young children who were emerging symbolic communicators on the autism spectrum in participating within a social routine. Results: Visual analysis and effect size estimates suggest the context-aware AAC condition resulted in increases in linguistic participation, vocal approximations, and visual attention for all three children. Conclusion: While this study was only an initial exploration and results are preliminary, context-aware AAC technologies have the potential to enhance participation and communication outcomes for young emerging symbolic communicators on the autism spectrum and more research is needed.

Journal Allard, J.B., Sharma, S., Patel, R., Sanderford, M., Tamura, K., Vucetic, S., Gerhard, G.S., Kumar, S., Evolutionary Sparse Learning Reveals the Shared Genetic Basis of Convergent Traits. Nature Communications, 16(1), p.3217, 2025.

Cases abound in which nearly identical traits have appeared in distant species facing similar environments. These unmistakable examples of adaptive evolution offer opportunities to gain insight into their genetic origins and mechanisms through comparative analyses. Here, we present an approach to build genetic models that underlie the independent origins of convergent traits using evolutionary sparse learning with paired species contrast (ESL-PSC). We tested the hypothesis that common genes and sites are involved in the convergent evolution of two key traits: C4 photosynthesis in grasses and echolocation in mammals. Genetic models were highly predictive of independent cases of convergent evolution of C4 photosynthesis. Genes contributing to genetic models for echolocation were highly enriched for functional categories related to hearing, sound perception, and deafness, a pattern that has eluded previous efforts applying standard molecular evolutionary approaches. These results support the involvement of sequence substitutions at common genetic loci in the evolution of convergent traits. Benchmarking on empirical and simulated datasets showed that ESL-PSC could be more sensitive in proteome-scale analyses to detect genes with convergent molecular evolution associated with the acquisition of convergent traits. We conclude that phylogeny-informed machine learning naturally excludes apparent molecular convergences due to shared species history, enhances the signal-to-noise ratio for detecting molecular convergence, and empowers the discovery of common genetic bases of trait convergences.

Journal Niknami, N., Mahzoon, V., Vucetic, S., Wu, J., Enhanced Meta-IDS: Adaptive Multi-Stage IDS with Sequential Model Adjustments. High-Confidence Computing, p.100298, 2025.

Traditional single-machine Network Intrusion Detection Systems (NIDS) are increasingly challenged by rapid network traffic growth and the complexities of advanced neural network methodologies. To address these issues, we propose an Enhanced Meta-IDS framework inspired by meta-computing principles, enabling dynamic resource allocation for optimized NIDS performance. Our hierarchical architecture employs a three-stage approach with iterative feedback mechanisms. We leverage these intervals in real-world scenarios with intermittent data batches to enhance our models. Outputs from the third stage provide labeled samples back to the first and second stages, allowing retraining and fine-tuning based on the most recent results without incurring additional latency. By dynamically adjusting model parameters and decision boundaries, our system optimizes responses to real-time data, effectively balancing computational efficiency and detection accuracy. By ensuring that only the most suspicious data points undergo intensive analysis, our multi-stage framework optimizes computational resource usage. Experiments on benchmark datasets demonstrate that our Enhanced Meta-IDS improves detection accuracy and reduces computational load or CPU time, ensuring robust performance in high-traffic environments. This adaptable approach offers an effective solution to modern network security challenges.

2024

Conference Russakoff, A., Miller, K., Mahzoon, V., Esmaeilkhani, P., Cho, C., Alzeidi, J., Hauri, S., Vucetic, S., CourtsightTV: An Interactive Visualization Software for Labeling Key Basketball Moments. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM), Boise, ID, 2024.

Advancements in sensor technology are leading to massive collection of tracking data in sports. There is an increasing interest in analyzing the tracking data to gain competitive advantage. Analyzing and labeling key game moments can provide deep insights into player performance, team dynamics, as well as game strategy. However, the process of manually labeling and analyzing these moments is costly and time-consuming. In this paper, we describe a visual interface for user-friendly and efficient labeling of key moments in basketball games aided by neural networks. We report results of a user study evaluating the labeling interface.

Conference Xu, H., Chen, M., Huang, L., Vucetic, S., Yin, W., X-Shot: A unified system to handle frequent, few-shot and zero-shot learning simultaneously in classification. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 4652–4665), 2024.

In recent years, few-shot and zero-shot learning, which learn to predict labels with limited annotated instances, have garnered significant attention. Traditional approaches often treat frequent-shot (freq-shot; labels with abundant instances), few-shot, and zero-shot learning as distinct challenges, optimizing systems for just one of these scenarios. Yet, in real-world settings, label occurrences vary greatly. Some of them might appear thousands of times, while others might only appear sporadically or not at all. For practical deployment, it is crucial that a system can adapt to any label occurrence. We introduce a novel classification challenge: X-shot, reflecting a real-world context where freq-shot, few-shot, and zero-shot labels co-occur without predefined limits. Here, X can span from 0 to positive infinity. The crux of X-shot centers on open-domain generalization and devising a system versatile enough to manage various label scenarios. To solve X-shot, we propose BinBin (Binary INference Based on INstruction following) that leverages the Indirect Supervision from a large collection of NLP tasks via instruction following, bolstered by Weak Supervision provided by large language models. BinBin surpasses previous state-of-the-art techniques on three benchmark datasets across multiple domains. To our knowledge, this is the first work addressing X-shot learning, where X remains variable.

Conference Yang, Z., Cherian, S., Vucetic, S., Two-pronged human evaluation of ChatGPT self-correction in radiology report simplification. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 4701–4714), 2024.

Radiology reports are highly technical documents aimed primarily at doctor-doctor communication. There has been an increasing interest in sharing those reports with patients, necessitating providing them patient-friendly simplifications of the original reports. This study explores the suitability of large language models in automatically generating those simplifications. We examine the usefulness of chain-of-thought and self-correction prompting mechanisms in this domain. We also propose a new evaluation protocol that employs radiologists and laypeople, where radiologists verify the factual correctness of simplifications, and laypeople assess simplicity and comprehension. Our experimental results demonstrate the effectiveness of self-correction prompting in producing high-quality simplifications. Our findings illuminate the preferences of radiologists and laypeople regarding text simplification, informing future research on this topic.

Conference Ara, Z., Ganguly, A., Peppard, D., Chung, D., Vucetic, S., Genaro Motti, V., Hong, S.R., Collaborative Job Seeking for People with Autism: Challenges and Design Opportunities. In Proceedings of the CHI Conference on Human Factors in Computing Systems (pp. 1-17), 2024.

Successful job search results from job seekers’ well-shaped social communication. While well-known diferences in communication exist between people with autism and neurotypicals, little is known about how people with autism collaborate with their social surroundings to strive in the job market. To better understand the practices and challenges of collaborative job seeking for people with autism, we interviewed 20 participants including applicants with autism, their social surroundings, and career experts. Through the interviews, we identified social challenges that people with autism face during their job seeking; the social support they leverage to be successful; and the technological limitations that hinder their collaboration. We designed four probes that represent major collaborative features found from the interviews–executive planning, communication, stage-wise preparation, and neurodivergent community formation–and discussed their potential usefulness and impact through three focus groups. We provide implications regarding how our findings can enhance collaborative job seeking experiences for people with autism through new designs.

Journal Lorah, E.R., MacNeil, S., Zimmerman, T., Rackensperger, T., Holyfield, C., Caldwell, N., Dragut, E.C., Vucetic, S., Spurring Innovation in AAC Technology through Collaborative Dreaming and Needs Finding with Individuals with Developmental Disabilities Who Use AAC. In Seminars in Speech and Language. Thieme Medical Publishers, Inc., 45(05), pp. 461-474, 2024.

Millions of individuals who have limited or no functional speech use augmentative and alternative communication (AAC) technology to participate in daily life and exercise the human right to communication. While advances in AAC technology lag significantly behind those in other technology sectors, mainstream technology innovations such as artificial intelligence (AI) present potential for the future of AAC. However, a new future of AAC will only be as effective as it is responsive to the needs and dreams of the people who rely upon it every day. AAC innovation must reflect an iterative, collaborative process with AAC users. To do this, we worked collaboratively with AAC users to complete participatory qualitative research about AAC innovation through AI. We interviewed 13 AAC users regarding (1) their current AAC engagement; (2) the barriers they experience in using AAC; (3) their dreams regarding future AAC development; and (4) reflections on potential AAC innovations. To analyze these data, a rapid research evaluation and appraisal was used. Within this article, the themes that emerged during interviews and their implications for future AAC development will be discussed. Strengths, barriers, and considerations for participatory design will also be described.

Journal Enayati, S., Vucetic, S., Leveraging Shortest Dependency Paths in Low-Resource Biomedical Relation Extraction, BMC Medical Informatics and Decision Making, 24(205), 2024.

Background: Biomedical Relation Extraction (RE) is essential for uncovering complex relationships between biomedical entities within text. However, training RE classifiers is challenging in low-resource biomedical applications with few labeled examples. Methods: We explore the potential of Shortest Dependency Paths (SDPs) to aid biomedical RE, especially in situations with limited labeled examples. In this study, we suggest various approaches to employ SDPs when creating word and sentence representations under supervised, semi-supervised, and in-context-learning settings. Results: Through experiments on three benchmark biomedical text datasets, we find that incorporating SDP-based representations enhances the performance of RE classifiers. The improvement is especially notable when working with small amounts of labeled data. Conclusion: SDPs offer valuable insights into the complex sentence structure found in many biomedical text passages. Our study introduces several straightforward techniques that, as demonstrated experimentally, effectively enhance the accuracy of RE classifiers.

Journal Holyfield, C., MacNeil, S., Caldwell, N., O'Neill Zimmerman, T., Lorah, E., Dragut, D., Vucetic, S., Leveraging Communication Partner Speech to Automate Augmented Input for Children on the Autism Spectrum Who Are Minimally Verbal: Prototype Development and Preliminary Efficacy Investigation, American Journal of Speech-Language Pathology, 33(3), 1174-1192, 2024.

Purpose: Augmentative and alternative communication (AAC) technology innovation is urgently needed to improve outcomes for children on the autism spectrum who are minimally verbal. One potential technology innovation is applying artificial intelligence (AI) to automate strategies such as augmented input to increase language learning opportunities while mitigating communication partner time and learning barriers. Innovation in AAC research and design methodology is also needed to empirically explore this and other applications of AI to AAC. The purpose of this report was to describe (a) the development of an AAC prototype using a design methodology new to AAC research and (b) a preliminary investigation of the efficacy of this potential new AAC capability. Method: The prototype was developed using a Wizard-of-Oz prototyping approach that allows for initial exploration of a new technology capability without the time and effort required for full-scale development. The preliminary investigation with three children on the autism spectrum who were minimally verbal used an adapted alternating treatment design to compare the effects of a Wizard-of-Oz prototype that provided automated augmented input (i.e., pairing color photos with speech) to a standard topic display (i.e., a grid display with line drawings) on visual attention, linguistic participation, and (for one participant) word learning during a circle activity. Results: Preliminary investigation results were variable, but overall participants increased visual attention and linguistic participation when using the prototype. Conclusions: Wizard-of-Oz prototyping could be a valuable approach to spur much needed innovation in AAC. Further research into efficacy, reliability, validity, and attitudes is required to more comprehensively evaluate the use of AI to automate augmented input in AAC.

2023

Conference Hauri S., Vucetic, S., Group Activity Recognition in Basketball Tracking Data--Neural Embeddings in Team Sports (NETS), In Proceedings of the 26th European Conference on Artificial Intelligence (ECAI), 2023.

Like many team sports, basketball involves two groups of players who engage in collaborative and adversarial activities to win a game. Players and teams are executing various complex strategies to gain an advantage over their opponents. Defining, identifying, and analyzing different types of activities is an important task in sports analytics, as it can lead to better strategies and decisions by the players and coaching staff. The objective of this paper is to automatically recognize basketball group activities from tracking data representing locations of players and the ball during a game. We propose a novel deep learning approach for group activity recognition (GAR) in team sports called NETS. To efficiently model the player relations in team sports, we combined a Transformer-based architecture with LSTM embedding, and a team-wise pooling layer to recognize the group activity. Training such a neural network generally requires a large amount of annotated data, which incurs high labeling cost. To alleviate this problem, we pretrain the neural network on a self-supervised trajectory prediction task and fine-tune it using a mix of strong and weak labels. We used a large tracking data set from 632 NBA games to evaluate our approach. The results show that NETS is capable of learning group activities with high accuracy, and that self- and weak-supervised training in NETS have a positive impact on GAR accuracy.

Conference Yang, Z., Cherian, S., Vucetic, S., Data Augmentation for Radiology Report Simplification. In Findings of the Association for Computational Linguistics: EACL 2023 (pp. 1877-1887), 2023.

This work considers the development of a text simplification model to help patients better understand their radiology reports. This paper proposes a data augmentation approach to address the data scarcity issue caused by the high cost of manual simplification. It prompts a large foundational pre-trained language model to generate simplifications of unlabeled radiology sentences. In addition, it uses paraphrasing of labeled radiology sentences. Experimental results show that the proposed data augmentation approach enables the training of a significantly more accurate simplification model than the baselines.

Journal Garrison, E., Singh, D., Hantula, D., Tincani, M., Nosek, J., Hong, S.R., Dragut, E., Vucetic, S., Understanding the Experience of Neurodivergent Workers in Image and Text Data Annotation. Computers in Human Behavior Reports, p.100318, 2023.

With the rise of large-scale data-driven innovation in AI, data annotation tasks found in digital work environments present an employment opportunity for neurodivergent individuals. Though work in data annotation can potentially ease the high unemployment rate of neurodivergent individuals, limited research focuses on the experience of neurodivergent workers in data annotation micro-tasks. To aid in understanding the experience of neurodivergent crowd workers, we conducted a user study with ten neurodivergent workers between the ages of 18–30. Participants completed three types of micro-tasks in a custom web-based data annotation platform. With the data collected from the platform, we examined individual responses within data annotations, work completion, the time to complete work, and calculated a potential “effective hourly wage,” for each participant based on their responses. Through a survey and semi-structured interview following each task, we learned about the experience of all participants regarding each of the data annotation tasks. Results of the study show: 1) our participants provide diverse annotations that are valuable for employers in digital data annotation work environments; 2) when calculating the “effective hourly wage” of all participants per task, most of our participants would earn less than minimum hourly wage on tasks; and 3) participant perceptions of the tasks matched their responses in the tasks presented.

Journal Maiti, A., Shi, S., Vucetic, S., An Ablation Study on the Use of Publication Venue Quality to Rank Computer Science Departments: Publication Quality is Strongly Correlated with the Subjective Perception of Research Strength. Scientometrics, pp.1-22, 2023.

This paper focuses on ranking computer science departments based on the quality of publications by the faculty in those departments. There are multiple strategies to convert publication lists into ranking scores for the departments. Important open questions include handling multi-author publications, inclusion criteria for publications and publication venues, accounting for the quality of publication venues, and accounting for the sub-areas of computer science. An ablation study is performed to evaluate the importance of different decisions for department ranking. The correlation between the resulting rankings and the peer assessment of computer science departments provided by the U.S. News was measured to evaluate the importance of different decisions. The results show that the selection of publication venues has the highest impact on the ranking. In contrast, decisions related to publication recency, multi-author publications, and clustering publications into subareas have less impact. Overall, Pearson’s correlation coefficient between the publication-based scores and the U.S. News ranking is above 0.90 for a large range of decisions, indicating a strong agreement between the objective measure and the subjective opinion of peers.

Journal Egleston, B.L., Bleicher, R.J., Fang, C.Y., Galloway, T.J., Vucetic, S., Benefits Versus Drawbacks of Delaying Surgery Due to Additional Consultations in Older Patients with Breast Cancer. Cancer Reports, 6(5), p.e1805, 2023.

Background: Additional evaluations, including second opinions, before breast cancer surgery may improve care, but may cause detrimental treatment delays that could allow disease progression. Aims We investigate the timing of surgical delays that are associated with survival benefits conferred by preoperative encounters versus the timing that are associated with potential harm. Methods and results We investigated survival outcomes of SEER Medicare patients with stage 1–3 breast cancer using propensity score‐based weighting. We examined interactions between the number of preoperative evaluation components and time from biopsy to definitive surgery. Components include new patient visits, unique surgeons, medical oncologists, or radiation oncologists consulted, established patient encounters, biopsies, and imaging studies. We identified 116 050 cases of whom 99% were female and had an average age of 75.0 ( SD = 6.2). We found that new patient visits have a protective association with respect to breast cancer mortality if they occur quickly after diagnosis with breast cancer mortality subdistribution Hazard Ratios [sHRs] = 0.87 (95% Confidence Interval [CI] 0.76–1.00) for 2, 0.71 (CI 0.55–0.92) for 3, and 0.63 (CI 0.37–1.07) for 4+ visits at minimal delay. New patient visits predict worsened mortality compared with no visits if the surgical delay is greater than 33 days (CI 14–53) for 2, 33 days (CI 17–49) for 3, and 44 days (CI 12–75) for 4+. Medical oncologist visits predict worse outcomes if the surgical delay is greater than 29 days (CI 20–39) for 1 and 38 days (CI 12–65) for 2+ visits. Similarly, surgeon encounters switch from a positive to a negative association if the surgical delay exceeds 29 days (CI 17–41) for 1 visit, but the positive estimate persists over time for 3+ surgeon visits. Conclusion: Preoperative visits that cause substantial delays may be associated with increased mortality in older patients with breast cancer.

Journal Tincani, M., Ji, H., Upthegrove, M., Garrison, E., West, M., Hantula, D., Vucetic, S., Dragut, E., Vocational Interventions for Individuals with ASD: Umbrella Review. Review Journal of Autism and Developmental Disorders, pp.1-37, 2023.

Adults with autism spectrum disorders (ASD) are disproportionately unemployed in comparison to adults without disabilities. In this umbrella review, we summarized the findings of 31 systematic reviews and meta-analyses of ASD vocational intervention research, encompassing 287 primary intervention studies. Most primary studies focused on strategies for teaching low-skill job tasks. In contrast, few primary studies examined interventions to improve more direct and meaningful measures of employment, such as job attainment, wages, hours worked, benefits, and job satisfaction. Over 60% of primary studies evaluated either behavioral or technology interventions, with a substantial overlap of primary studies between reviews. Few primary studies explored other approaches, including comprehensive support strategies, accommodations typically found in the workplace, and strategies for teaching job seeking. Intervention effects were generally positive; however, review study quality was variable and low for many reviews, constraining the interpretation of findings. Most reviews reported incomplete primary study participant demographic information, rendering it challenging to generalize findings outside of the reviews. Our results highlight the need for a greater focus on outcome measures indicative of actual employment, the evaluation of a greater variety of interventions, including accommodations typically found in the workplace, and the need to assess comprehensive support. Low review quality underscores the need for improved methodological rigor in reviews of vocational intervention research.

Journal Egleston, B.L., Chanda, A.K., Bai, T., Fang, C.Y., Bleicher, R.J., Vucetic, S., Using Pointwise Mutual Information for Breast Cancer Health Disparities Research with SEER-Medicare Claims. Methodology: European journal of research methods for the behavioral & social sciences, 19(1), p.43, 2023.

Identification of procedures using International Classification of Diseases or Healthcare Common Procedure Coding System codes is challenging when conducting medical claims research. We demonstrate how Pointwise Mutual Information can be used to find associated codes. We apply the method to an investigation of racial differences in breast cancer outcomes. We used Surveillance Epidemiology and End Results (SEER) data linked to Medicare claims. We identified treatment using two methods. First, we used previously published definitions. Second, we augmented definitions using codes empirically identified by the Pointwise Mutual Information statistic. Similar to previous findings, we found that presentation differences between Black and White women closed much of the estimated survival curve gap. However, we found that survival disparities were completely eliminated with the augmented treatment definitions. We were able to control for a wider range of treatment patterns that might affect survival differences between Black and White women with breast cancer.

2022

Conference Xu, H., Vucetic, S., Yin, W., OpenStance: Real-world Zero-shot Stance Detection. In Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL) (pp. 314-324), 2022.

Prior studies of zero-shot stance detection identify the attitude of texts towards unseen topics occurring in the same document corpus. Such task formulation has three limitations: (i) Single domain/dataset. A system is optimized on a particular dataset from a single domain; therefore, the resulting system cannot work well on other datasets; (ii) the model is evaluated on a limited number of unseen topics; (iii) it is assumed that part of the topics has rich annotations, which might be impossible in real-world applications. These drawbacks will lead to an impractical stance detection system that fails to generalize to open domains and open-form topics. This work defines OpenStance: open-domain zero-shot stance detection, aiming to handle stance detection in an open world with neither domain constraints nor topic-specific annotations. The key challenge of OpenStance lies in open-domain generalization: learning a system with fully unspecific supervision but capable of generalizing to any dataset. To solve OpenStance, we propose to combine indirect supervision, from textual entailment datasets, and weak supervision, from data generated automatically by pre-trained Language Models. Our single system, without any topic-specific supervision, outperforms the supervised method on three popular datasets. To our knowledge, this is the first work that studies stance detection under the open-domain zero-shot setting. All data and code will be publicly released.

Conference Chanda, A., Egleston, B., Bai, T., Vucetic, S., MedCV: An Interactive Visualization System for Patient Cohor tIdentification from Medical Claim Data, Proceedings of the 31th ACM International Conference on Information & Knowledge Management (CIKM), Atlanta, GA, 2022.

Healthcare providers generate a medical claim after every patient visit. A medical claim consists of a list of medical codes describing the diagnosis and any treatment provided during the visit. Medical claims have been popular in medical research as a data source for retrospective cohort studies. This paper introduces a medical claim visualization system (MedCV) that supports cohort selection from medical claim data. MedCV was developed as part of a design study in collaboration with clinical researchers and statisticians. It helps a researcher to define inclusion rules for cohort selection by revealing relationships between medical codes and visualizing medical claims and patient timelines. Evaluation of our system through a user study indicates that MedCV enables domain experts to define high-quality inclusion rules in a time-efficient manner.

Conference Enayati, S., Vucetic, S., MERIT: Minimal Supervision Through Label Augmentation for Biomedical Relation Extraction, Proceedings of the American Medical Informatics Association Symposium (AMIA), Washington D.C., 2022.

Relation Extraction (RE) is an important task in extracting structured data from free biomedical text. Obtaining labeled data needed to train RE models in specialized domains such as biomedicine can be very expensive because it requires expert ...

Journal Lindner, G., Shi, S., Vucetic, S., Miskovic, S., Transfer Learning for Radioactive Particle Tracking, Chemical Engineering Science, 248, p.117190, 2022.

Radioactive particle tracking (RPT) is a non-invasive technique used to monitor opaque multiphase flow systems. Achieving highly accurate particle tracing is challenging and time-consuming because of the need to build a new RPT model from calibration data each time the experimental conditions change. This paper aims to examine if RPT calibration data under previous conditions can be leveraged with the help of transfer learning (TL) when creating an RPT model for a new condition. Several TL strategies for exploiting historical calibration data are evaluated in conjunction with Geant4 simulations to understand their applicability to RPT. The results show that when it is impractical to collect a lot of calibration data, TL is often superior to training an RPT model only on new data. Moreover, when new calibration data collection is not feasible, an RPT model trained on historical data can be very accurate if sufficiently similar to the historical conditions.

Journal Chanda, A.K., Bai, T., Yang, Z., Vucetic, S., Improving Medical Term Embeddings Using UMLS Metathesaurus, BMC Medical Informatics and Decision Making, 22(1), pp.1-12, 2022.

Background: Health providers create Electronic Health Records (EHRs) to describe the conditions and procedures used to treat their patients. Medical notes entered by medical staff in the form of free text are a particularly insightful component of EHRs. There is a great interest in applying machine learning tools on medical notes in numerous medical informatics applications. Learning vector representations, or embeddings, of terms in the notes, is an important pre-processing step in such applications. However, learning good embeddings is challenging because medical notes are rich in specialized terminology, and the number of available EHRs in practical applications is often very small. Methods: In this paper, we propose a novel algorithm to learn embeddings of medical terms from a limited set of medical notes. The algorithm, called definition2vec , exploits external information in the form of medical term definitions. It is an extension of a skip-gram algorithm that incorporates textual definitions of medical terms provided by the Unified Medical Language System (UMLS) Metathesaurus. Results: To evaluate the proposed approach, we used a publicly available Medical Information Mart for Intensive Care (MIMIC-III) EHR data set. We performed quantitative and qualitative experiments to measure the usefulness of the learned embeddings. The experimental results show that definition2vec keeps the semantically similar medical terms together in the embedding vector space even when they are rare or unobserved in the corpus. We also demonstrate that learned vector embeddings are helpful in downstream medical informatics applications. Conclusion: This paper shows that medical term definitions can be helpful when learning embeddings of rare or previously unseen medical terms from a small corpus of specialized documents such as medical notes.

Journal Wiese, D., Lynch, S.M., Stroup, A.M., Maiti, A., Harris, G., Vucetic, S., Henry, K.A., Examining Socio-Spatial Mobility Patterns Among Colon Cancer Patients After Diagnosis, SSM-Population Health, 17, p.101023, 2022.

Given the growing number of cancer survivors, it is important to better understand socio-spatial mobility patterns of cancer patients after diagnosis that could have public health implications regarding post-diagnostic access to care for treatment and follow-up surveillance. In this exploratory study, residential histories from LexisNexis were linked to New Jersey colon cancer cases diagnosed from 2006 to 2011 to examine differences in socio-spatial mobility patterns after diagnosis by stage at cancer diagnosis, sex, and race/ethnicity. For the colon cancer cases, we summarized and compared the number of residences and changes in the residential census tract and neighborhood poverty after the diagnosis. We found only minor changes in neighborhood poverty among the cases during the follow-up period after diagnosis. During the follow-up period of up to 10 years after diagnosis, 67% of the patients did not move to a different residential census tract, and 10.8% moved from New Jersey to another state. Cases that moved to a different census tract changed after diagnosis were generally less wealthy than non-movers, but the destination of relocation varied by race/ethnicity and socioeconomic status. We also found a significant association between residential mobility and stage at diagnosis, whereby patients diagnosed with colon cancer at an early stage were more likely to be movers. This study contributes to understanding of the socio-spatial mobility patterns in colon cancer patients and may help to inform cancer research by summarizing the extent to which colon cancer patients move after diagnosis.

2021

Conference Gao, H., Xu, H., Vucetic, S., Sample Efficient Decentralized Stochastic Frank-Wolfe Methods for Continuous DR-Submodular Maximization, 30th International Joint Conference on Artificial Intelligence (IJCAI), Montreal, Canada, 2021.

Continuous DR-submodular maximization is an important machine learning problem, which covers numerous popular applications. With the emergence of large-scale distributed data, developing efficient algorithms for the continuous DR-submodular maximization, such as the decentralized Frank-Wolfe method, became an important challenge. However, existing decentralized Frank-Wolfe methods for this kind of problem have the sample complexity of O(1/ε³), incurring a large computational overhead. In this paper, we propose two novel sample efficient decentralized Frank-Wolfe methods to address this challenge. Our theoretical results demonstrate that the sample complexity of the two proposed methods is O(1/ε²), which is better than O(1/ε³) of the existing methods. As far as we know, this is the first published result achieving such a favorable sample complexity. Extensive experimental results confirm the effectiveness of the proposed methods.

Conference Enayati, S., Yang, Z., Lu, B., Vucetic, S., A Visualization Approach for Rapid Labeling of Clinical Notes for Smoking Status Extraction, Proceedings of the Second Workshop on Data Science with Human in the Loop: Language Advances, 2021.

Labeling is typically the most human-intensive step during the development of supervised learning models. In this paper, we propose a simple and easy-to-implement visualization approach that reduces cognitive load and increases the speed of text labeling. The approach is fine-tuned for task of extraction of patient smoking status from clinical notes. The proposed approach consists of the ordering of sentences that mention smoking, centering them at smoking tokens, and annotating to enhance informative parts of the text. Our experiments on clinical notes from the MIMIC-III clinical database demonstrate that our visualization approach enables human annotators to label sentences up to 3 times faster than with a baseline approach.

Conference He, L., Shen, C., Mukherjee, A., Vucetic, S., Dragut, E., Cannot Predict Comment Volume of a News Article before (a few) Users Read It, 15th International Conference on Web and Social Media (ICWSM), 2021.

Many news outlets allow users to contribute comments on topics about daily world events. News articles are the seeds that spring users' interest to contribute content, i.e., comments. An article may attract an apathetic user engagement (several tens of comments) or a spontaneous fervent user engagement (thousands of comments). In this paper, we study the problem of predicting the total number of user comments a news article will receive. Our main insight is that user-to-user interaction factors contribute the most to an accurate prediction, while news article specific factors have surprisingly little influence. This appears to be an interesting and understudied phenomenon: collective social behavior at a news outlet shapes user response and may even downplay the content of an article. We compile and analyze a large number of features, both old and novel from literature. The features span a broad spectrum of facets including news article and comment contents, temporal dynamics, sentiment/linguistic features, and user behaviors. We show that the early arrival rate of comments is the best indicator of the eventual number of comments. We conduct an in-depth analysis of this feature across several dimensions, such as news outlets and news article categories. We show that the relationship between the early rate and the final number of comments as well as the prediction accuracy vary considerably across news outlets and news article categories (e.g., politics, sports, or health).

Conference Hauri, S., Djuric, N., Radosavljevic, V., Vucetic, S., Multi-Modal Trajectory Prediction of NBA Players, IEEE Winter Conference on Applications of Computer Vision (WACV), 2021.

National Basketball Association (NBA) players are highly motivated and skilled experts that solve complex decision making problems at every time point during a game. As a step towards understanding how players make their decisions, we focus on their movement trajectories during games. We propose a method that captures the multi-modal behavior of players, where they might consider multiple trajectories and select the most advantageous one. The method is built on an LSTM-based architecture predicting multiple trajectories and their probabilities, trained by a multi-modal loss function that updates the best trajectories. Experiments on large, fine-grained NBA tracking data show that the proposed method outperforms the state-of-the-art. In addition, the results indicate that the approach generates more realistic trajectories and that it can learn individual playing styles of specific players.

Journal McGee, F., Novinger, Q., Hauri, S., Vucetic, S., Levy, R.M., Carnevale, V., Haldane, A., Generative Capacity of Probabilistic Protein Sequence Models, Nature Communications, 12(1), p. 1-14, 2021.

Potts models and variational autoencoders (VAEs) have recently gained popularity as generative protein sequence models (GPSMs) to explore fitness landscapes and predict the effect of mutations. Despite encouraging results, quantitative characterization and comparison of GPSM-generated probability distributions is still lacking. It is currently unclear whether GPSMs can faithfully reproduce the complex multi-residue mutation patterns observed in natural sequences arising due to epistasis. We develop a set of sequence statistics to comparatively assess the accuracy, or “generative capacity”, of three GPSMs: a pairwise Potts Hamiltonian, a vanilla VAE, and a site-independent model, using natural and synthetic datasets. We show that the generative capacity of the Potts Hamiltonian model is the largest; the higher order mutational statistics generated by the model agree with those observed for natural sequences. In contrast, we show that the vanilla VAE’s generative capacity lies between the pairwise Potts and site-independent models. Importantly, our work measures GPSM generative capacity in terms of higher-order sequence covariation and provides a new framework for evaluating and interpreting GPSM accuracy that emphasizes the role of epistasis.

Journal Banjade, H.R., Hauri, S., Zhang, S., Ricci, F., Gong, W., Hautier, G., Vucetic, S., Yan, Q., Structure Motif–Centric Learning Framework for Inorganic Crystalline Systems, Science Advances, 7(17), p.eabf1754, 2021.

Incorporation of physical principles in a machine learning (ML) architecture is a fundamental step toward the continued development of artificial intelligence for inorganic materials. As inspired by the Pauling’s rule, we propose that structure motifs in inorganic crystals can serve as a central input to a machine learning framework. We demonstrated that the presence of structure motifs and their connections in a large set of crystalline compounds can be converted into unique vector representations using an unsupervised learning algorithm. To demonstrate the use of structure motif information, a motif-centric learning framework is created by combining motif information with the atom-based graph neural networks to form an atom-motif dual graph network (AMDNet), which is more accurate in predicting the electronic structures of metal oxides such as bandgaps. The work illustrates the route toward fundamental design of graph neural network learning architecture for complex materials by incorporating beyond-atom physical principles.

Journal Henry, K.A., Wiese, D., Maiti, A., Harris, G., Vucetic, S., Stroup, A.M., Geographic Clustering of Cutaneous T-Cell Lymphoma in New Jersey: an Exploratory Analysis using Residential Histories, Cancer Causes & Control, pp.1-11, 2021.

Cutaneous T-cell lymphoma (CTCL) is a rare type of non-Hodgkin lymphoma. Previous studies have reported geographic clustering of CTCL based on the residence at the time of diagnosis. We explore geographic clustering of CTCL using both the residence at the time of diagnosis and past residences using data from the New Jersey State Cancer Registry. CTCL cases (n = 1,163) diagnosed between 2006–2014 were matched to colon cancer controls (n = 17,049) on sex, age, race/ethnicity, and birth year. Jacquez's Q-Statistic was used to identify temporal clustering of cases compared to controls. Geographic clustering was assessed using the Bernoulli-based scan-statistic to compare cases to controls, and the Poisson-based scan-statisic to compare the observed number of cases to the number expected based on the general population. Significant clusters (p < 0.05) were mapped, and standard incidence ratios (SIR) reported. We adjusted for diagnosis year, sex, and age. The Q-statistic identified significant temporal clustering of cases based on past residences in the study area from 1992 to 2002. A cluster was detected in 1992 in Bergen County in northern New Jersey based on the Bernoulli (1992 SIR 1.84) and Poisson (1992 SIR 1.86) scan-statistics. Using the Poisson scan-statistic with the diagnosis location, we found evidence of an elevated risk in this same area, but the results were not statistically significant. There is evidence of geographic clustering of CTCL cases in New Jersey based on past residences. Additional studies are necessary to understand the possible reasons for the excess of CTCL cases living in this specific area some 8–14 years prior to diagnosis.

Journal Lachaud, A., Marcus, A., Vucetic, S., Miskovic, I., Study of the Influence of On-Deposit Locations in Data-Driven Mineral Prospectivity Mapping: A Case Study on the Iskut Project in Northwestern British Columbia, Canada, Minerals, 11(6), p.597, 2021.

The accuracy of data-driven predictive mineral prospectivity models relies heavily on the training datasets used. These models are usually trained using data for “known” deposit locations as well as “non-deposit” locations that are based on randomly generated point patterns. In this study, data related to the Seabridge Gold Inc Iskut project, an epithermal Au deposit in northwestern British Columbia (BC), Canada, are used to test the utility of data-driven mineral prospectivity modeling. The input spatial dataset is comprised mostly of publicly available data. Data for 18 vein and epithermal Au known mineral occurrences (KMO) are obtained from the BC Geological Survey’s MINFILE repository and selected as training deposit locations. A total of eleven sets of non-deposit locations (NDL) were also created, including one set of selected non-prospective KMO for Au deposits from the MINFILE and ten sets of random point patterns. Given the scale of this study, most of the KMO recorded on the property are of the epithermal deposit type. Hence, they could not be used as a selection criterion. Data-driven mineral potential models are generated using the random forest (RF) algorithm and trained on multiple data sets. The comparison of RF models demonstrated that using non-prospective KMO generates more accurate predictions than the random point pattern. The produced mineral prospectivity maps delineated multiple areas with higher discovery potential, which matched viable targets for the Au-Cu epithermal-porphyry system identified through previous Seabridge Gold Inc. (Toronto, ON, Canada) field reconnaissance and drilling programs.

Journal Wiese, D., Stroup, A.M., Maiti, A., Harris, G., Lynch, S.M., Vucetic, S., Gutierrez-Velez, V.H., Henry, K.A., Measuring Neighborhood Landscapes: Associations between a Neighborhood’s Landscape Characteristics and Colon Cancer Survival, International Journal of Environmental Research and Public Health, 18(9), p.4728, 2021.

Landscape characteristics have been shown to influence health outcomes, but few studies have examined their relationship with cancer survival. We used data from the National Land Cover Database to examine associations between regional-stage colon cancer survival and 27 different landscape metrics. The study population included all adult New Jersey residents diagnosed between 2006 and 2011. Cases were followed until 31 December 2016 (N = 3949). Patient data were derived from the New Jersey State Cancer Registry and were linked to LexisNexis to obtain residential histories. Cox proportional hazard regression was used to estimate hazard ratios (HR) and 95% confidence intervals (CI95) for the different landscape metrics. An increasing proportion of high-intensity developed lands with 80–100% impervious surfaces per cell/pixel was significantly associated with the risk of colon cancer death (HR = 1.006; CI95 = 1.002–1.01) after controlling for neighborhood poverty and other individual-level factors. In contrast, an increase in the aggregation and connectivity of vegetation-dominated low-intensity developed lands with 20–<40% impervious surfaces per cell/pixel was significantly associated with the decrease in risk of death from colon cancer (HR = 0.996; CI95 = 0.992–0.999). Reducing impervious surfaces in residential areas may increase the aesthetic value and provide conditions more advantageous to a healthy lifestyle, such as walking. Further research is needed to understand how these landscape characteristics impact survival.

2020

Conference Ha, P., Zhang, S., Djuric, N., Vucetic, S., Improving Word Embeddings through Iterative Refinement of Word- and Character-level Models, 28th International Conference on Computational Linguistics (COLING), Barcelona, Spain, 2020. (first author undergrad student)

Embedding of rare and out-of-vocabulary (OOV) words is an important open NLP problem. A popular solution is to train a character-level neural network to reproduce the embeddings from a standard word embedding model. The trained network is then used to assign vectors to any input string, including OOV and rare words. We enhance this approach and introduce an algorithm that iteratively refines and improves both word- and character-level models. We demonstrate that our method outperforms the existing algorithms on 5 word similarity data sets, and that it can be successfully applied to job title normalization, an important problem in the e-recruitment domain that suffers from the OOV problem.

Conference Djuric, N., Wang, Z., Vucetic, S., Growing Adaptive Multi-Hyperplane Machines, 37th International Conference on Machine Learning (ICML), Vienna, Austria, 2020.

Adaptive Multi-hyperplane Machine (AMM) is an online algorithm for learning Multi-hyperplane Machine (MM), a classification model which allows multiple hyperplanes per class. AMM is based on Stochastic Gradient Descent (SGD), with training time comparable to linear Support Vector Machine (SVM) and significantly higher accuracy. On the other hand, empirical results indicate there is a large accuracy gap between AMM and non-linear SVMs. In this paper we show that this performance gap is not due to limited representability of the MM model, as it can represent arbitrary concepts. We set to explain the connection between the AMM and Learning Vector Quantization (LVQ) algorithms, and introduce a novel Growing AMM (GAMM) classifier motivated by Growing LVQ, that imputes duplicate hyperplanes into the MM model during SGD training. We provide theoretical results showing that GAMM has favorable convergence properties, and analyze the generalization bound of the MM models. Experiments indicate that GAMM achieves significantly improved accuracy on non-linear problems, with only slightly slower training compared to AMM. On some tasks GAMM comes close to non-linear SVM, and outperforms other popular classifiers such as Neural Networks and Random Forests.

Journal Wiese, D., Stroup, A.M., Maiti, A., Harris, G., Lynch, S.M., Vucetic, S., Henry, K.A., Residential Mobility and Geospatial Disparities in Colon Cancer Survival. Cancer Epidemiology and Prevention Biomarkers, 29(11), pp.2119-2125, 2020.

Background: Identifying geospatial cancer survival disparities is critical to focus interventions and prioritize efforts with limited resources. Incorporating residential mobility into spatial models may result in different geographic patterns of survival compared with the standard approach using a single location based on the patient's residence at the time of diagnosis. Methods: Data on 3,949 regional-stage colon cancer cases diagnosed from 2006 to 2011 and followed until December 31, 2016, were obtained from the New Jersey State Cancer Registry. Geographic disparity based on the spatial variance and effect sizes from a Bayesian spatial model using residence at diagnosis was compared with a time-varying spatial model using residential histories [adjusted for sex, gender, substage, race/ethnicity, and census tract (CT) poverty]. Geographic estimates of risk of colon cancer death were mapped. Results: Most patients (65%) remained at the same residence, 22% changed CT, and 12% moved out of state. The time-varying model produced a wider range of adjusted risk of colon cancer death (0.85–1.20 vs. 0.94–1.11) and resulted in greater geographic disparity statewide after adjustment (25.5% vs. 14.2%) compared with the model with only the residence at diagnosis. Conclusions: Including residential mobility may allow for more precise estimates of spatial risk of death. Results based on the traditional approach using only residence at diagnosis were not substantially different for regional stage colon cancer in New Jersey. Impact: Including residential histories opens up new avenues of inquiry to better understand the complex relationships between people and places, and the effect of residential mobility on cancer outcomes.

Journal Egleston, B.L., Bai, T., Bleicher, R.J., Taylor, S.J., Lutz, M.H., Vucetic, S., Statistical Inference for Natural Language Processing Algorithms with a Demonstration Using Type 2 Diabetes Prediction from Electronic Health Record Notes. Biometrics, published online July 22, 2020.

The pointwise mutual information statistic (PMI), which measures how often two words occur together in a document corpus, is a cornerstone of recently proposed popular natural language processing algorithms such as word2vec. PMI and word2vec reveal semantic relationships between words and can be helpful in a range of applications such as document indexing, topic analysis, or document categorization. We use probability theory to demonstrate the relationship between PMI and word2vec. We use the theoretical results to demonstrate how the PMI can be modeled and estimated in a simple and straight forward manner. We further describe how one can obtain standard error estimates that account for within‐patient clustering that arises from patterns of repeated words within a patient's health record due to a unique health history. We then demonstrate the usefulness of PMI on the problem of predictive identification of disease from free text notes of electronic health records. Specifically, we use our methods to distinguish those with and without type 2 diabetes mellitus in electronic health record free text data using over 400 000 clinical notes from an academic medical center.

Journal Wiese, D., Stroup, A.M., Maiti, A., Harris, G., Lynch, S.M., Vucetic, S., Henry, K.A., Socioeconomic Disparities in Colon Cancer Survival: Revisiting Neighborhood Poverty using Residential Histories. Epidemiology, 31(5), pp.728-735, 2020.

Background: Residential histories linked to cancer registry data provide new opportunities to examine cancer outcomes by neighborhood socioeconomic status (SES). We examined differences in regional stage colon cancer survival estimates comparing models using a single neighborhood SES at diagnosis to models using neighborhood SES from residential histories. Methods: We linked regional stage colon cancers from the New Jersey State Cancer Registry diagnosed from 2006 to 2011 to LexisNexis administrative data to obtain residential histories. We defined neighborhood SES as census tract poverty based on location at diagnosis and across the follow-up period through 31 December 2016 based on residential histories (average, time-weighted average, time-varying). Using Cox proportional hazards regression, we estimated associations between colon cancer and census tract poverty measurements (continuous and categorical), adjusted for age, sex, race/ethnicity, regional substage, and mover status. Results: Sixty-five percent of the sample was nonmovers (one census tract); 35% (movers) changed tract at least once. Cases from tracts with >20% poverty changed residential tracts more often (42%) than cases from tracts with <5% poverty (32%). Hazard ratios (HRs) were generally similar in strength and direction across census tract poverty measurements. In time-varying models, cases in the highest poverty category (>20%) had a 30% higher risk of regional stage colon cancer death than cases in the lowest category (<5%) (95% confidence interval [CI] = 1.04, 1.63). Conclusion: Residential changes after regional stage colon cancer diagnosis may be associated with a higher risk of colon cancer death among cases in high-poverty areas. This has important implications for postdiagnostic access to care for treatment and follow-up surveillance.

Journal Shapovalov, M., Dunbrack Jr, R.L., Vucetic, S., Multifaceted Analysis of Training and Testing Convolutional Neural Networks for Protein Secondary Structure Prediction. PloS One, 15(5), p.e0232528, 2020.

Protein secondary structure prediction remains a vital topic with broad applications. Due to lack of a widely accepted standard in secondary structure predictor evaluation, a fair comparison of predictors is challenging. A detailed examination of factors that contribute to higher accuracy is also lacking. In this paper, we present: (1) new test sets, Test2018, Test2019, and Test2018-2019, consisting of proteins from structures released in 2018 and 2019 with less than 25% identity to any protein published before 2018; (2) a 4-layer convolutional neural network, SecNet, with an input window of ±14 amino acids which was trained on proteins ≤25% identical to proteins in Test2018 and the commonly used CB513 test set; (3) an additional test set that shares no homologous domains with the training set proteins, according to the Evolutionary Classification of Proteins (ECOD) database; (4) a detailed ablation study where we reverse one algorithmic choice at a time in SecNet and evaluate the effect on the prediction accuracy; (5) new 4- and 5-label prediction alphabets that may be more practical for tertiary structure prediction methods. The 3-label accuracy (helix, sheet, coil) of the leading predictors on both Test2018 and CB513 is 81-82%, while SecNet's accuracy is 84% for both sets. Accuracy on the non-homologous ECOD set is only 0.6 points (83.9%) lower than the results on the Test2018-2019 set (84.5%). The ablation study of features, neural network architecture, and training hyper-parameters suggests the best accuracy results are achieved with good choices for each of them while the neural network architecture is not as critical as long as it is not too simple. Protocols for generating and using unbiased test, validation, and training sets are provided. Our data sets, including input features and assigned labels, and SecNet software including third-party dependencies and databases, are downloadable from dunbrack.fccc.edu/ss and github.com/sh-maxim/ss.

2019

Conference Sondur, S, Kant, K, Vucetic, S., Storage on the Edge: Evaluating Cloud Backed Edge Storage in Cyberphysical Systems, 16th International Conference on Mobile Ad-hoc and Smart Systems (MASS), Monterey, CA, 2019.

Effective control of emerging cyberphysical systems such as smart transportation, smart health-care, etc. requires edge computing infrastructure that is often organized into three layers, namely edge (IoT) devices, edge controllers (ECs) and the cloud. In large infrastructures, ECs must be deployed densely in the proximity of edge devices and need to satisfy strict constraints on cost, size, cooling, etc. Thus, ECs cannot host large amounts of local storage and instead must make use of cloud storage in the background to provide an impression of large, fast local storage to host the IoT device data needed for online and real-time queries. In this paper, we provide insights into the configuration issues of such an edge storage infrastructure (ESI) based on the evaluation of commercial ESIs on several real-world edge computing workloads. We also show that the current ESI designs are lacking in several respects, and suggest some approaches for enhancing their capabilities to meet the stringent requirements of emerging edge computing applications.

Conference Zhang, S., He, L., Dragut, E., Vucetic, S., How to Invest my Time: Lessons from Human-in- the-Loop Entity Extraction, 25th ACM SIGKDD International Conf. on Knowledge Discovery and Data Mining (KDD), Anchorage, AK, 2019.

Recognizing entities that follow or closely resemble a regular expression (regex) pattern is an important task in information extraction. Common approaches for extraction of such entities require humans to either write a regex recognizing an entity or manually label entity mentions in a document corpus. While human effort is critical to build an entity recognition model, surprisingly little is known about how to best invest that effort given a limited time budget. To get an answer, we consider an iterative human-in-the-loop (HIL) framework that allows users to write a regex or manually label entity mentions, followed by training and refining a classifier based on the provided information. We demonstrate on 5 entity recognition tasks that classification accuracy improves over time with either approach. When a user is allowed to choose between regex construction and manual labeling, we discover that (1) if the time budget is low, spending all time for regex construction is often advantageous, (2) if the time budget is high, spending all time for manual labeling seems to be superior, and (3) between those two extremes, writing regexes followed by manual labeling is typically the best approach. Our code and data is available at https://github.com/nymph332088/HILRecognizer.

Conference Maiti, A., Vucetic, S., Spatial Aggregation Facilitates Discovery of Spatial Topics, 57th Annual Meeting of the Association for Computational Linguistics (ACL), Florence, Italy, 2019.

Spatial aggregation refers to merging of documents created at the same spatial location. We show that by spatial aggregation of a large collection of documents and applying a traditional topic discovery algorithm on the aggregated data we can efficiently discover spatially distinct topics. By looking at topic discovery through matrix factorization lenses we show that spatial aggregation allows low rank approximation of the original document-word matrix, in which spatially distinct topics are preserved and non-spatial topics are aggregated into a single topic. Our experiments on synthetic data confirm this observation. Our experiments on 4.7 million tweets collected during the Sandy Hurricane in 2012 show that spatial and temporal aggregation allows rapid discovery of relevant spatial and temporal topics during that period. Our work indicates that different forms of document aggregation might be effective in rapid discovery of various types of distinct topics from large collections of documents.

Conference Bai, T., Egleston, B.L., Bleicher, R., Vucetic, S., Medical Concept Representation Learning from Multi-Source Data, 28th International Joint Conf. on Artificial Intelligence (IJCAI), Macao, China, 2019.

Representing words as low dimensional vectors is very useful in many natural language processing tasks. This idea has been extended to medical domain where medical codes listed in medical claims are represented as vectors to facilitate exploratory analysis and predictive modeling. However, depending on a type of a medical provider, medical claims can use medical codes from different ontologies or from a combination of ontologies, which complicates learning of the representations. To be able to properly utilize such multi-source medical claim data, we propose an approach that represents medical codes from different ontologies in the same vector space. We first modify the Pointwise Mutual Information (PMI) measure of similarity between the codes. We then develop a new negative sampling method for word2vec model that implicitly factorizes the modified PMI matrix. The new approach was evaluated on the code cross-reference problem, which aims at identifying similar codes across different ontologies. In our experiments, we evaluated cross-referencing between ICD-9 and CPT medical code ontologies. Our results indicate that vector representations of codes learned by the proposed approach provide superior cross-referencing when compared to several existing approaches.

Conference Bai, T., Vucetic, S., Improving Medical Code Prediction from Clinical Text via Incorporating Online Knowledge Sources, The World Wide Web Conference (WWW), 72-82, San Francisco, CA, 2019.

Clinical notes contain detailed information about health status of patients for each of their encounters with a health system. Developing effective models to automatically assign medical codes to clinical notes has been a long-standing active research area. Despite a great recent progress in medical informatics fueled by deep learning, it is still a challenge to find the specific piece of evidence in a clinical note which justifies a particular medical code out of all possible codes. Considering the large amount of online disease knowledge sources, which contain detailed information about signs and symptoms of different diseases, their risk factors, and epidemiology, there is an opportunity to exploit such sources. In this paper we consider Wikipedia as an external knowledge source and propose Knowledge Source Integration (KSI), a novel end-to-end code assignment framework, which can integrate external knowledge during training of any baseline deep learning model. The main idea of KSI is to calculate matching scores between a clinical note and disease related Wikipedia documents, and combine the scores with output of the baseline model. To evaluate KSI, we experimented with automatic assignment of ICD-9 diagnosis codes to the emergency department clinical notes from MIMIC-III data set, aided by Wikipedia documents corresponding to the ICD-9 codes. We evaluated several baseline models, ranging from logistic regression to recently proposed deep learning models known to achieve the state-of-the-art accuracy on clinical notes. The results show that KSI consistently improves the baseline models and that it is particularly successful in assignment of rare codes. In addition, by analyzing weights of KSI models, we can gain understanding about which words in Wikipedia documents provide useful information for predictions.

Journal Zhou, N., Jiang, Y., … Vucetic, S., … Friedberg, I., The CAFA Challenge Reports Improved Protein Function Prediction and New Functional Annotations for Hundreds of Genes Through Experimental Screens. Genome Biology, 20(1), pp.1-23, 2019.

Background: The Critical Assessment of Functional Annotation (CAFA) is an ongoing, global, community-driven effort to evaluate and improve the computational annotation of protein function. Results: Here, we report on the results of the third CAFA challenge, CAFA3, that featured an expanded analysis over the previous CAFA rounds, both in terms of volume of data analyzed and the types of analysis performed. In a novel and major new development, computational predictions and assessment goals drove some of the experimental assays, resulting in new functional annotations for more than 1000 genes. Specifically, we performed experimental whole-genome mutation screening in Candida albicans and Pseudomonas aureginosa genomes, which provided us with genome-wide experimental data for genes associated with biofilm formation and motility. We further performed targeted assays on selected genes in Drosophila melanogaster , which we suspected of being involved in long-term memory. Conclusion: We conclude that while predictions of the molecular function and biological process annotations have slightly improved over time, those of the cellular component have not. Term-centric prediction of experimental annotations remains equally challenging; although the performance of the top methods is significantly better than the expectations set by baseline methods in C. albicans and D. melanogaster , it leaves considerable room and need for improvement. Finally, we report that the CAFA community now involves a broad range of participants with expertise in bioinformatics, biological experimentation, biocuration, and bio-ontologies, working together to improve functional annotation, computational function prediction, and our ability to manage big data in the era of large experimental screens.

Journal Shapovalov, M., Vucetic, S., Dunbrack, R.L., A New Clustering and Nomenclature for Beta Turns Derived from High-Resolution Protein Structures, PLoS Computational Biology, 15 (3), 2019.

Protein loops connect regular secondary structures and contain 4-residue beta turns which represent 63% of the residues in loops. The commonly used classification of beta turns (Type I, I', II, II', VIa1, VIa2, VIb, and VIII) was developed in the 1970s and 1980s from analysis of a small number of proteins of average resolution, and represents only two thirds of beta turns observed in proteins (with a generic class Type IV representing the rest). We present a new clustering of beta-turn conformations from a set of 13,030 turns from 1074 ultra-high resolution protein structures (≤1.2 Å). Our clustering is derived from applying the DBSCAN and k-medoids algorithms to this data set with a metric commonly used in directional statistics applied to the set of dihedral angles from the second and third residues of each turn. We define 18 turn types compared to the 8 classical turn types in common use. We propose a new 2-letter nomenclature for all 18 beta-turn types using Ramachandran region names for the two central residues (e.g., 'A' and 'D' for alpha regions on the left side of the Ramachandran map and 'a' and 'd' for equivalent regions on the right-hand side; classical Type I turns are 'AD' turns and Type I' turns are 'ad'). We identify 11 new types of beta turn, 5 of which are sub-types of classical beta-turn types. Up-to-date statistics, probability densities of conformations, and sequence profiles of beta turns in loops were collected and analyzed. A library of turn types, BetaTurnLib18, and cross-platform software, BetaTurnTool18, which identifies turns in an input protein structure, are freely available and redistributable from dunbrack.fccc.edu/betaturn and github.com/sh-maxim/BetaTurn18. Given the ubiquitous nature of beta turns, this comprehensive study updates understanding of beta turns and should also provide useful tools for protein structure determination, refinement, and prediction programs.

2018

Conference Zhang, S., Pal, A., Kant, K., Vucetic, S., Enhancing Disaster Situational Awareness via Automated Summary Dissemination of Social Media Content, 2018 IEEE Global Communications Conference (GLOBECOM), Abu Dhabi, UEA, 2018.

The paper proposes a situational awareness service, named StayTuned, that collects information from social media, extracts relevant messages, and broadcasts them to the subscribers through wireless emergency alert system. StayTuned uses automated filtering and summarization of messages and updates subscribers with real-time situational summaries. Extensive experiments were conducted using Twitter data collected during the Sandy hurricane to evaluate performance of the automated message extraction.

Conference Zhang, S., He, L., Vucetic, S., Dragut, E., Regular Expression Guided Entity Mention Mining from Noisy Web Data, Conference on Empirical Methods in Natural Language Processing (EMNLP), Brussels, Belgium, 2018.

Many important entity types in web documents, such as dates, times, email addresses, and course numbers, follow or closely resemble patterns that can be described by Regular Expressions (REs). Due to a vast diversity of web documents and ways in which they are being generated, even seemingly straightforward tasks such as identifying mentions of date in a document become very challenging. It is reasonable to claim that it is impossible to create a RE that is capable of identifying such entities from web documents with perfect precision and recall. Rather than abandoning REs as a go-to approach for entity detection, this paper explores ways to combine the expressive power of REs, ability of deep learning to learn from large data, and human-in-the loop approach into a new integrated framework for entity identification from web data. The framework starts by creating or collecting the existing REs for a particular type of an entity. Those REs are then used over a large document corpus to collect weak labels for the entity mentions and a neural network is trained to predict those RE-generated weak labels. Finally, a human expert is asked to label a small set of documents and the neural network is fine tuned on those documents. The experimental evaluation on several entity identification problems shows that the proposed framework achieves impressive accuracy, while requiring very modest human effort.

Conference Bai, T., Zhang, S., Egleston, B.L., Vucetic, S., Interpretable Representation Learning for Healthcare via Capturing Disease Progression through Time, 24th ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining (KDD), London, UK, 2018.

Various deep learning models have recently been applied to predictive modeling of Electronic Health Records (EHR). In medical claims data, which is a particular type of EHR data, each patient is represented as a sequence of temporally ordered irregularly sampled visits to health providers, where each visit is recorded as an unordered set of medical codes specifying patient’s diagnosis and treatment provided during the visit. Based on the observation that different patient conditions have different temporal progression patterns, in this paper we propose a novel interpretable deep learning model, called Timeline. The main novelty of Timeline is that it has a mechanism that learns time decay factors for every medical code. This allows the Timeline to learn that chronic conditions have a longer lasting impact on future visits than acute conditions. Timeline also has an attention mechanism that improves vector embeddings of visits. By analyzing the attention weights and disease progression functions of Timeline, it is possible to interpret the predictions and understand how risks of future visits change over time. We evaluated Timeline on two large-scale real world data sets. The specific task was to predict what is the primary diagnosis category for the next hospital visit given previous visits. Our results show that Timeline has higher accuracy than the state of the art deep learning models based on RNN. In addition, we demonstrate that time decay factors and attentions learned by Timeline are in accord with the medical knowledge and that Timeline can provide a useful insight into its predictions.

Journal Vucetic, S., Chanda, A.K., Zhang, S., Bai, T., Maiti, A., Faculty Citation Measures are Highly Correlated with Peer Assessment of Computer Science Doctoral Programs, Communications of the ACM, 2018.

We study relationship between peer assessment of quality of U.S. Computer Science (CS) doctoral programs and objective measures of research strength of those programs. In Fall 2016 we collected Google Scholar citation data for 4,352 tenure-track CS faculty from 173 U.S. universities. The citations are measured by the t10 index, which represents the number of citations received by the 10th highest cited paper of a faculty. To measure the research strength of a CS doctoral program we use 2 groups of citation measures. The first group of measures averages t10 of faculty in a program. Pearson correlation of those measures with the peer assessment of U.S. CS doctoral programs published by the U.S. News in 2014 is as high as 0.890. The second group of measures counts the number of well cited faculty in a program. Pearson correlation of those measures with the peer assessment is as high as 0.909. By combining those two groups of measures using linear regression, we create the Scholar score whose Pearson correlation with the peer assessment is 0.933 and which explains 87.2% of the variance in the peer assessment. Our evaluation shows that the highest 62 ranked CS doctoral programs by the U.S. News peer assessment are much higher correlated with the Scholar score than the next 57 ranked programs, indicating the deficiencies of peer assessment of less-known CS programs. Our results also indicate that university reputation might have a sizeable impact on peer assessment of CS doctoral programs. To promote transparency, the raw data and the codes used in this study are made available to research community at http://www.dabi.temple.edu/~vucetic/CSranking/.

Journal Bai, T., Chanda, A.K., Egleston, B.L., Vucetic, S., EHR Phenotyping via Jointly Embedding Medical Concepts and Words into a Unified Vector Space, BMC Medical Informatics and Decision Making, 2018.

BACKGROUND: There has been an increasing interest in learning low-dimensional vector representations of medical concepts from Electronic Health Records (EHRs). Vector representations of medical concepts facilitate exploratory analysis and predictive modeling of EHR data to gain insights about the patterns of care and health outcomes. EHRs contain structured data such as diagnostic codes and laboratory tests, as well as unstructured free text data in form of clinical notes, which provide more detail about condition and treatment of patients. METHODS: In this work, we propose a method that jointly learns vector representations of medical concepts and words. This is achieved by a novel learning scheme based on the word2vec model. Our model learns those relationships by integrating clinical notes and sets of accompanying medical codes and by defining joint contexts for each observed word and medical code. RESULTS: In our experiments, we learned joint representations using MIMIC-III data. Using the learned representations of words and medical codes, we evaluated phenotypes for 6 diseases discovered by our and baseline method. The experimental results show that for each of the 6 diseases our method finds highly relevant words. We also show that our representations can be very useful when predicting the reason for the next visit. CONCLUSIONS: The jointly learned representations of medical concepts and words capture not only similarity between codes or words themselves, but also similarity between codes and words. They can be used to extract phenotypes of different diseases. The representations learned by the joint model are also useful for construction of patient features.

2016

Conference Zhang, S., Vucetic, S., Sampling Bias in LinkedIn: A Case Study, 25th International World Wide Web Conference (WWW), Montreal, Canada, 2016.

This paper describes a case study of sampling bias in LinkedIn, a major professional social network. The study collected a sample of 1,989 STEM students who graduated from a major public university between 2002 and 2014. Overall, 40% of the graduates had a LinkedIn profile in summer of 2015. It was observed that LinkedIn participation significantly fluctuated among different majors, and ranged from 30% for biochemistry majors to 51% for information science majors. Year of graduation, gender, and grade point average surprisingly did not seem to create a large difference in LinkedIn participation. These results should be useful for design and interpretation of empirical studies which use LinkedIn data or select participants from LinkedIn social network.

Conference Han, C., Zhang, S., Ghalwash, M., Vucetic, S., Obradovic, Z., Joint Learning of Representation and Structure for Sparse Regression on Graphs, 16th SIAM Conference on Data Mining (SDM), Miami, FL, 2016.

In many applications, including climate science, power systems, and remote sensing, multiple input variables are observed for each output variable and the output variables are dependent. Several methods have been proposed to improve prediction by learning the conditional distribution of the output variables. However, when the relationship between the raw features and the outputs is nonlinear, the existing methods cannot capture both the nonlinearity and the underlying structure well. In this study, we propose a structured model containing hidden variables, which are nonlinear functions of inputs and which are linearly related with the output variables. The parameters modeling the relationships between the input and hidden variables, between the hidden and output variables, as well as among the output variables are learned simultaneously. To demonstrate the effectiveness of our proposed method, we conducted extensive experiments on eight synthetic datasets and three real-world challenging datasets: forecasting wind power, forecasting solar energy, and forecasting precipitation over U.S. The proposed method was more accurate than state-of-the-art structured regression methods.

Conference Djuric, N., Grbovic, M., Vucetic, S., ParkAssistant: An Algorithm for Guiding a Car to a Parking Spot, Transportation Research Board 95th Annual Meeting (TRB), Washington, D.C., USA, 2016.

Parking search is a major issue in urban areas. Drivers in major cities face a daily struggle in finding parking space, and much of this is due to lack of information about parking rules, parking prices, traffic conditions, and parking availability. As a consequence, drivers often perform inefficient search for a parking space and spend too much time searching, pay too much, or park too far from an intended destination. Inefficient parking search is not a problem only for drivers, it is also increasing traffic congestion and pollution and it causes a distortion of the parking market. Despite the vast technological advances in recent decades, parking search remains fundamentally the same societal problem it has been for almost a century. The objective of this paper is to address this issue by proposing the ParkAssistant, an algorithm that calculates a cruising route that minimizes the expected cost of parking, defined as a mix of price and time to reach the destination. To calculate a good cruising route, the algorithm uses parking information that consists of parking rules, traffic conditions, probabilities of finding an empty parking space, and drivers’ utility function. We evaluated ParkAssistant through simulations using real-life parking occupancy data from San Francisco, CA. The results indicate that, as compared to an uninformed driver model, it allows drivers to find parking much faster. The results also show that the quality of ParkAssistant recommendations grows with the quality of parking information the method is provided with.

Journal Jiang, Y., et al. An Expanded Evaluation of Protein Function Prediction Methods Shows an Improvement in Accuracy, Genome Biology, Vol. 17(1), 184, 2016.

BACKGROUND: A major bottleneck in our understanding of the molecular underpinnings of life is the assignment of function to proteins. While molecular experiments provide the most reliable annotation of proteins, their relatively low throughput and restricted purview have led to an increasing role for computational function prediction. However, assessing methods for protein function prediction and tracking progress in the field remain challenging. RESULTS: We conducted the second critical assessment of functional annotation (CAFA), a timed challenge to assess computational methods that automatically assign protein function. We evaluated 126 methods from 56 research groups for their ability to predict biological functions using Gene Ontology and gene-disease associations using Human Phenotype Ontology on a set of 3681 proteins from 18 species. CAFA2 featured expanded analysis compared with CAFA1, with regards to data set size, variety, and assessment metrics. To review progress in the field, the analysis compared the best methods from CAFA1 to those of CAFA2. CONCLUSIONS: The top-performing methods in CAFA2 outperformed those from CAFA1. This increased accuracy can be attributed to a combination of the growing number of experimental annotations and improved methods for function prediction. The assessment also revealed that the definition of top-performing algorithms is ontology specific, that different performance metrics can be used to probe the nature of accurate predictions, and the relative diversity of predictions in the biological process and human phenotype ontologies. While there was methodological improvement between CAFA1 and CAFA2, the interpretation of results and usefulness of individual methods remain context-dependent.

Journal Djuric, N., Kansakar, L., Vucetic, S., Semi-Supervised Combination of Experts for Aerosol Optical Depth Estimation, Artificial Intelligence Journal, Vol. 230, pp. 1-13, 2016.

Combination of experts Aerosol estimation Remote sensing Aerosols are small airborne particles produced by natural and man-made sources. Aerosol Optical Depth (AOD), recognized as one of the most important quantities in understanding and predicting the Earth’s climate, is estimated daily on a global scale by several Earthobserving satellite instruments. Each instrument has different coverage and sensitivity to atmospheric and surface conditions, and, as a result, the quality of AOD estimated by different instruments varies across the globe. We present a semi-supervised method for learning how to aggregate estimations from multiple satellite instruments into a more accurate estimate, where labels come from a small number of accurate and expensive ground-based instruments. The method also accounts for the problem of missing experts, an issue inherent to the AOD estimation task. By assuming a context-dependent prior, the model is capable of incorporating additional information and providing estimates even when there are no available experts. Moreover, the proposed method uses a latent variable to partition the data, so that in each partition the expert AOD estimations are aggregated in a different, optimal way. We applied the method to combine global AOD estimations from 5 instruments aboard 4 satellites, and the results indicate it can successfully exploit labeled and unlabeled data to produce accurate aggregated AOD estimations.

2015

Journal Zhang, K., Lan, L., Kwok, T.J., Vucetic, S., Parvim, B., Large Scale Semi-Supervised Learning via Sparse Nonparametric Prototype Model, IEEE Transactions on Neural Networks and Learning Systems, Vol. 26 (3), pp. 444-457, 2015.

When the amount of labeled data are limited, semisupervised learning can improve the learner’s performance by also using the often easily available unlabeled data. In particular, a popular approach requires the learned function to be smooth on the underlying data manifold. By approximating this manifold as a weighted graph, such graph-based techniques can often achieve state-of-the-art performance. However, their high time and space complexities make them less attractive on large data sets. In this paper, we propose to scale up graph-based semisupervised learning using a set of sparse prototypes derived from the data. These prototypes serve as a small set of data representatives, which can be used to approximate the graph-based regularizer and to control model complexity. Consequently, both training and testing become much more efficient. Moreover, when the Gaussian kernel is used to define the graph affinity, a simple and principled method to select the prototypes can be obtained. Experiments on a number of real-world data sets demonstrate encouraging performance and scaling properties of the proposed approach. It also compares favorably with models learned via 1-regularization at the same level of model sparsity. These results demonstrate the efficacy of the proposed approach in producing highly parsimonious and accurate models for semisupervised learning.

2014

Conference Radosavljevic, V., Vucetic, S., Obradovic, Z., Neural Gaussian Conditional Random Fields, Machine Learning and Knowledge Discovery in Databases (ECML), Nancy, France, 2014.

We propose a Conditional Random Field (CRF) model for structured regression. By constraining the feature functions as quadratic functions of outputs, the model can be conveniently represented in a Gaussian canonical form. We improved the representational power of the resulting Gaussian CRF (GCRF) model by (1) introducing an adaptive feature function that can learn nonlinear relationships between inputs and outputs and (2) allowing the weights of feature functions to be dependent on inputs. Since both the adaptive feature functions and weights can be constructed using feedforward neural networks, we call the resulting model Neural GCRF. The appeal of Neural GCRF is in conceptual simplicity and computational efficiency of learning and inference through use of sparse matrix computations. Experimental evaluation on the remote sensing problem of aerosol estimation from satellite measurements and on the problem of document retrieval showed that Neural GCRF is more accurate than the benchmark predictors.

Conference Lan, L., Malbasa, V., Vucetic, S., Spatial Scan for Disease Mapping on a Mobile Population, 28th AAAI Conference on Artificial Intelligence (AAAI), Quebec City, Canada, 2014.

In disease mapping, the spatial scan statistic is used to detect spatial regions where population is exposed to a significantly higher disease risk than expected. In this important application, the current residence is typically used to define the location of individuals from the population. Considering the mobility of humans at various temporal and spatial scales, using only information about the current residence may be an insufficiently informative proxy because it ignores a multitude of exposures that may occur away from home, or which had occurred at previous residences. In this paper, we propose a spatial scan statistic that is appropriate for disease mapping on mobile populations. We formulate a computationally efficient algorithm that uses the proposed statistic to find significant high-risk regions from mobile population's disease status data. The algorithm is applicable on large populations and over dense spatial grids. The experimental results demonstrate that the proposed algorithm is computationally efficient and outperforms the traditional disease clustering approaches at discovering high-risk regions in mobile populations.

Conference Djuric, N., Grbovic, M., Radosavljevic, V., Bhamidipati, N., Vucetic, S., Non-linear Label Ranking for Large-scale Prediction of Long-Term User Interests, 28th AAAI Conference on Artificial Intelligence (AAAI), Quebec City, Canada, 2014.

We consider the problem of personalization of online services from the viewpoint of ad targeting, where we seek to find the best ad categories to be shown to each user, resulting in improved user experience and increased advertiser's revenue. We propose to address this problem as a task of ranking the ad categories depending on a user's preference, and introduce a novel label ranking approach capable of efficiently learning non-linear, highly accurate models in large-scale settings. Experiments on real-world advertising data set with more than 3.2 million users show that the proposed algorithm outperforms the existing solutions in terms of both rank loss and top-K retrieval performance, strongly suggesting the benefit of using the proposed model on large-scale ranking problems.

Conference Coric, V., Djuric, N., Vucetic, S., Frugal Traffic Monitoring with Autonomous Participatory Sensing, SIAM Conference on Data Mining (SDM), Philadelphia, PA, 2014.

As mobile devices are becoming pervasive, participatory sensing is becoming an attractive way of collecting large quantities of valuable location-based data. An important participatory sensing application is traffic monitoring, where GPS-enabled smartphones can provide invaluable information about traffic conditions. In this paper we propose a strategy for frugal sensing in which the participants send only a fraction of the observed traffic information to reduce costs while achieving high accuracy. The strategy is based on autonomous sensing, in which participants make decisions to send traffic information without guidance from the central server, thus reducing the communication overhead and improving privacy. We propose to use traffic flow theory in deciding whether or not to send an observation to the server. To provide accurate and computationally efficient estimation of the current traffic, we propose to use a budgeted version of the Gaussian Process model on the server side. The model is tuned to be robust to missing observations, which occur quite often in frugal sensing. The experiments on a real-life traffic data set indicate that the proposed approach can use up to two orders of magnitude fewer samples than a baseline approach when estimating traffic speed on a highway network, with only a negligible loss in accuracy.

Conference Grbovic M., Vucetic S., Generating Ad Targeting Rules using Sparse Principal Component Analysis with Constraints, International World Wide Web Conference (WWW), 2014.

Determining the right audience for an advertising campaign is a well-established problem, of central importance to many Internet companies. Two distinct targeting approaches exist, the model-based approach, which leverages machine learning, and the rule-based approach, which relies on manual generation of targeting rules. Common rules include identifying users that had interactions (website visits, emails received, etc.) with the companies related to the advertiser, or search queries related to their product. We consider a problem of discovering such rules from data using Constrained Sparse PCA. The constraints are put in place to account for cases when evidence in data suggests a relation that is not appropriate for advertising. Experiments on real-world data indicate the potential of the proposed approach.

Journal Djuric, N., Radosavljevic, V., Obradovic, Z., Vucetic, S., Gaussian Conditional Random Fields for Aggregation of Operational Aerosol Retrievals, IEEE Geoscience and Remote Sensing Letters, Vol. 12, no. 4, pp. 761-765, 2014.

We present a Gaussian conditional random field model for the aggregation of aerosol optical depth (AOD) retrievals from multiple satellite instruments into a joint retrieval. The model provides aggregated retrievals with higher accuracy and coverage than any of the individual instruments while also providing an estimation of retrieval uncertainty. The proposed model finds an optimal temporally smoothed combination of individual retrievals that minimizes the root-mean-squared error of AOD retrieval. We evaluated the model on five years (2006–2010) of satellite data over North America from five instruments (Aqua and Terra MODIS, MISR, SeaWiFS, and the Ozone Monitoring Instrument), collocated with ground-based Aerosol Robotic Network ground-truth AOD readings, clearly showing that the aggregation of different sources leads to improvements in the accuracy and coverage of AOD retrievals.

2013

Conference Djuric, N., Vucetic, S., Efficient Visualization of Large-scale Data Tables through Reordering and Entropy Minimization, IEEE International Conference on Data Mining (ICDM), Dallas, TX, 2013.

Visualization of data tables with n examples and m columns using heatmaps provides a holistic view of the original data. As there are n! ways to order rows and m! ways to order columns, and data tables are typically ordered without regard to visual inspection, heatmaps of the original data tables often appear as noisy images. However, if rows and columns of a data table are ordered such that similar rows and similar columns are grouped together, a heatmap may provide a deep insight into the underlying data distribution. We propose an information-theoretic approach to produce a well-ordered data table. In particular, we search for ordering that minimizes entropy of residuals of predictive coding applied on the ordered data table. This formalization leads to a novel ordering procedure, EM-ordering, that can be applied separately on rows and columns. For ordering of rows, EM-ordering repeats until convergence the steps of rescaling columns and solving a Traveling Salesman Problem where rows are treated as cities. To allow fast ordering of large data tables, we propose an efficient TSP heuristic with modest O(n log n) time complexity. When compared to existing state-of-the-art reordering approaches, the method often provides heatmaps of higher visual quality while being significantly more scalable. Analysis of real-world traffic and financial data sets further confirms that EM-ordering can be a valuable tool for visual exploration of large-scale data sets.

Conference Djuric, N., Kansakar, L., Vucetic, S., Semi-Supervised Learning for Integration of Aerosol Predictions from Multiple Satellite Instruments, 23rd International Joint Conference on Artificial Intelligence (IJCAI), Beijing, China, 2013. (Outstanding IJCAI paper: the best paper in the AI and Computational Sustainability Track)

Aerosol Optical Depth (AOD), recognized as one of the most important quantities in understanding and predicting the Earth’s climate, is estimated daily on a global scale by several Earth-observing satellite instruments. Each instrument has different coverage and sensitivity to atmospheric and surface conditions, and, as a result, the quality of AOD estimated by different instruments varies across the globe. We present a method for learning how to aggregate AOD estimations from multiple satellite instruments into a more accurate estimation. The proposed method is semi-supervised, as it is able to learn from a small number of labeled data, where labels come from a few accurate and expensive ground-based instruments, and a large number of unlabeleddata. Themethodusesalatentvariableto partitionthedata, sothatineachpartitiontheexpert AOD estimations are aggregated in a different, optimalway. Weapplied themethodtocombineAOD estimations from 5 instruments aboard 4 satellites, and the results indicate that it can successfully exploitlabeledandunlabeleddatatoproduceaccurate aggregated AOD estimations.

Conference Grbovic, M., Djuric, N., Vucetic, S., Multi-prototype Label Ranking with Novel Pairwise to Total Rank Aggregation, 23rd International Joint Conference on Artificial Intelligence (IJCAI), Beijing, China, 2013.

We propose a multi-prototype-based algorithm for onlinelearningofsoftpairwise-preferencesoverlabels. The algorithm learns soft label preferences via minimization of the proposed soft rank-loss measure, and can learn from total orders as well as from various types of partial orders. The soft pairwise preference algorithm outputs are further aggregated to produce a total label ranking prediction using a novel aggregation algorithm that outperforms existing aggregation solutions. Experiments on synthetic and real-world data demonstrate stateof-the-art performance of the proposed model.

Conference Ristovski, K., Radosavljevic, V., Vucetic, S., Obradovic, Z., Continuous Conditional Random Fields for Efficient Regression in Large Fully Connected Graphs, 27th AAAI Conference on Artificial Intelligence (AAAI), Bellevue, WA, 2013.

When used for structured regression, powerful Conditional Random Fields (CRFs) are typically restricted to modeling effects of interactions among examples in local neighborhoods. Using more expressive representation would result in dense graphs, making these methods impractical for large-scale applications. To address this issue, we propose an effective CRF model with linear scale-up properties regarding approximate learning and inference for structured regression on large, fully connected graphs. The proposed method is validated on real-world large-scale problems of image de-noising and remote sensing. In conducted experiments, we demonstrated that dense connectivity provides an improvement in prediction accuracy. Inference time of less than ten seconds on graphs with millions of nodes and trillions of edges makes the proposed model an attractive tool for large-scale, structured regression problems.

Journal Djuric, N., Lan, L., Vucetic, S., Wang, Z., BudgetedSVM: A Toolbox for Scalable SVM Approximations, Journal of Machine Learning Research, Vol. 14, pp.3813−3817, 2013.

We present BudgetedSVM, an open-source C++ toolbox comprising highly-optimized implementations of recently proposed algorithms for scalable training of Support Vector Machine (SVM) approximators: Adaptive Multi-hyperplane Machines, Low-rank Linearization SVM, and Budgeted Stochastic Gradient Descent. BudgetedSVM trains models with accuracy comparable to LibSVM in time comparable to LibLinear, solving non-linear problems with millions of high-dimensional examples within minutes on a regular computer. We provide command-line and Matlab interfaces to BudgetedSVM, an efficient API for handling large-scale, high-dimensional data sets, as well as detailed documentation to help developers use and further extend the toolbox.

Journal Grbovic M., Li W., Subrahmanya N. A., Usadi A. K., Vucetic, S., Cold Start Approach for Data Driven Fault Detection, IEEE Transactions on Industrial Informatics, Vol. 9(4), pp. 2264 – 2273, 2013.

A typical assumption in supervised fault detection is that abundant historical data are available prior to model learning, where all types of faults have already been observed at least once. This assumption is likely to be violated in practical settings as new fault types can emerge over time. In this paper we study this often overlooked cold start learning problem in data-driven fault detection, where in the beginning only normal operation data are available and faulty operation data become available as the faults occur. We explored how to leverage strengths of unsupervised and supervised approaches to build a model capable of detecting faults even if none are still observed, and of improving over time, as new fault types are observed. The proposed framework was evaluated on the benchmark Tennessee Eastman Process data. The proposed fusion model performed better on both unseen and seen faults than the stand-alone unsupervised and supervised models.

Journal Lan, L., Vucetic, S., Multi-task Feature Selection in Microarray Data by Binary Integer Programming, BMC Bioinformatics, Vol. 7(Suppl 7):S5, 2013.

A major challenge in microarray classification is that the number of features is typically orders of magnitude larger than the number of examples. In this paper, we propose a novel feature filter algorithm to select the feature subset with maximal discriminative power and minimal redundancy by solving a quadratic objective function with binary integer constraints. To improve the computational efficiency, the binary integer constraints are relaxed and a low-rank approximation to the quadratic term is applied. The proposed feature selection algorithm was extended to solve multi-task microarray classification problems. We compared the single-task version of the proposed feature selection algorithm with 9 existing feature selection methods on 4 benchmark microarray data sets. The empirical results show that the proposed method achieved the most accurate predictions overall. We also evaluated the multi-task version of the proposed algorithm on 8 multi-task microarray datasets. The multi-task feature selection algorithm resulted in significantly higher accuracy than when using the single-task feature selection methods.

Journal Grbovic, M., Djuric, N., Guo, S., Vucetic, S., Supervised Clustering of Label Ranking Data using Label Preference Information, Machine Learning Journal, Vol. 93 (2-3), pp 191-225, 2013.

This paper studies supervised clustering in the context of label ranking data. The goal is to partition the feature space into K clusters, such that they are compact in both the feature and label ranking space. This type of clustering has many potential applications. For example, in target marketing we might want to come up with K different offers or marketing strategies for our target audience. Thus, we aim at clustering the customers’ feature space into K clusters by leveraging the revealed or stated, potentially incomplete customer preferences over products, such that the preferences of customers within one cluster are more similar to each other than to those of customers in other clusters. We establish several baseline algorithms and propose two principled algorithms for supervised clustering. In the first baseline, the clusters are created in an unsupervised manner, followed by assigning a representative label ranking to each cluster. In the second baseline, the label ranking space is clustered first, followed by partitioning the feature space based on the central rankings. In the third baseline, clustering is applied on a new feature space consisting of both features and label rankings, followed by mapping back to the original feature and ranking space. The RankTree principled approach is based on a Ranking Tree algorithm previously proposed for label ranking prediction. Our modification starts with K random label rankings and iteratively splits the feature space to minimize the ranking loss, followed by re-calculation of the K rankings based on cluster assignments. The MM-PL approach is a Editors: Eyke Hüllermeier and Johannes Fürnkranz. M. Grbovic ( ) · N. Djuric · S. Vucetic Department of Computer and Information Sciences, Center for Data Analytics and Biomedical Informatics, Temple University, Philadelphia, PA 19122, USA e-mail: mihajlo.grbovic@temple.edu N. Djuric e-mail: nemanja.djuric@temple.edu S. Vucetic e-mail: slobodan.vucetic@temple.edu S. Guo Xerox Research Centre Europe, 6 chemin de Maupertuis, 38240 Meylan, France e-mail: shengbo.guo@xrce.xerox.com 192 Mach Learn (2013) 93:191–225 multi-prototype supervised clustering algorithm based on the Plackett-Luce (PL) probabilistic ranking model. It represents each cluster with a union of Voronoi cells that are defined by a set of prototypes, and assign each cluster with a set of PL label scores that determine the cluster central ranking. Cluster membership and ranking prediction for a new instance are determined by cluster membership of its nearest prototype. The unknown cluster PL parameters and prototype positions are learned by minimizing the ranking loss, based on two variants of the expectation-maximization algorithm. Evaluation of the proposed algorithms was conducted on synthetic and real-life label ranking data by considering several measures of cluster goodness: (1) cluster compactness in feature space, (2) cluster compactness in label ranking space and (3) label ranking prediction loss. Experimental results demonstrate that the proposed MM-PL and RankTree models are superior to the baseline models. Further, MM-PL is has shown to be much better than other algorithms at handling situations with significant fraction of missing label preferences.

Journal Grbovic M., Vucetic S., Decentralized Estimation using Distortion Sensitive Learning Vector Quantization, Pattern Recognition Letters, Vol. 34 (9), pp. 963–969, 2013.

A typical approach in supervised learning when data comes from multiple sources is to send original data from all sources to a central location and train a predictor that estimates a certain target quantity. This can be inefficient and costly in applications with constrained communication channels, due to limited power and/or bitlength constraints. Under such constraints, one potential solution is to send encoded data from sources and use a decoder at the central location. Data at each source is summarized into a single codeword and sent to a central location, where a target quantity is estimated using received codewords. This problem is known as Decentralized Estimation. In this paper we propose a variant of the Learning Vector Quantization (LVQ) classification algorithm, the Distortion Sensitive LVQ (DSLVQ), to be used for encoder design in decentralized estimation. Unlike most related research that assumes known distributions of source observations, we assume that only a set of empirical samples is available. DSLVQ approach is compared to previously proposed Regression Tree and Deterministic Annealing (DA) approaches for encoder design in the same setting. While Regression Tree is very fast to train, it is limited to encoder regions with axis-parallel splits. On the other hand, DA is known to provide state-of-the-art performance. However, its training complexity grows with the number of sources that have different data distributions, due to over-parametrization. Our experiments on several synthetic and one real-world remote sensing problem show that DA has limited application potential as it is highly impractical to train even in a four-source setting, while DSLVQ is as simple and fast to train as the Regression Tree. In addition, DSLVQ shows similar performance to DA in experiments with small number of sources and outperforms DA in experiments with large number of sources, while consistently outperforming the Regression Tree algorithm.

Journal Lan, L., Djuric, N., Guo, Y., Vucetic, S., MS-kNN: Protein Function Prediction by Integrating Multiple Data Sources, BMC Bioinformatics, Vol. 14 (suppl. 3):S8, 2013.

BACKGROUND: Protein function determination is a key challenge in the post-genomic era. Experimental determination of protein functions is accurate, but time-consuming and resource-intensive. A cost-effective alternative is to use the known information about sequence, structure, and functional properties of genes and proteins to predict functions using statistical methods. In this paper, we describe the Multi-Source k-Nearest Neighbor (MS-kNN) algorithm for function prediction, which finds k-nearest neighbors of a query protein based on different types of similarity measures and predicts its function by weighted averaging of its neighbors' functions. Specifically, we used 3 data sources to calculate the similarity scores: sequence similarity, protein-protein interactions, and gene expressions. RESULTS: We report the results in the context of 2011 Critical Assessment of Function Annotation (CAFA). Prior to CAFA submission deadline, we evaluated our algorithm on 1,302 human test proteins that were represented in all 3 data sources. Using only the sequence similarity information, MS-kNN had term-based Area Under the Curve (AUC) accuracy of Gene Ontology (GO) molecular function predictions of 0.728 when 7,412 human training proteins were used, and 0.819 when 35,622 training proteins from multiple eukaryotic and prokaryotic organisms were used. By aggregating predictions from all three sources, the AUC was further improved to 0.848. Similar result was observed on prediction of GO biological processes. Testing on 595 proteins that were annotated after the CAFA submission deadline showed that overall MS-kNN accuracy was higher than that of baseline algorithms Gotcha and BLAST, which were based solely on sequence similarity information. Since only 10 of the 595 proteins were represented by all 3 data sources, and 66 by two data sources, the difference between 3-source and one-source MS-kNN was rather small. CONCLUSIONS: Based on our results, we have several useful insights: (1) the k-nearest neighbor algorithm is an efficient and effective model for protein function prediction; (2) it is beneficial to transfer functions across a wide range of organisms; (3) it is helpful to integrate multiple sources of protein information.

Journal Radivojac, P., Clark, W. T., ..., Toppo, S., Lan, L., Djuric, N., Guo, Y., Vucetic, S., Bairoch, A., Linial, M., Babbitt, P. C., et al., A Large-scale Evaluation of Computational Protein Function Prediction, Nature Methods, Vol. 10(3), pp. 221-229, 2013.

Automated annotation of protein function is challenging. As the number of sequenced genomes rapidly grows, the overwhelming majority of protein products can only be annotated computationally. If computational predictions are to be relied upon, it is crucial that the accuracy of these methods be high. Here we report the results from the first large-scale community-based critical assessment of protein function annotation (CAFA) experiment. Fifty-four methods representing the state of the art for protein function prediction were evaluated on a target set of 866 proteins from 11 organisms. Two findings stand out: (i) today's best protein function prediction algorithms substantially outperform widely used first-generation methods, with large gains on all types of targets; and (ii) although the top methods perform well enough to guide experiments, there is considerable need for improvement of currently available tools.

2012

Conference Grbovic M., Dance. C., Vucetic S., Sparse Principal Component Analysis with Constraints, 26th AAAI Conference on Artificial Intelligence (AAAI), Toronto, CA, 2012.

The sparse principal component analysis is a variant of the classical principal component analysis, which finds linear combinations of a small number of features that maximize variance across data. In this paper we propose a methodology for adding two general types of feature grouping constraints into the original sparse PCA optimization procedure.We derive convex relaxations of the considered constraints, ensuring the convexity of the resulting optimization problem. Empirical evaluation on three real-world problems, one in process monitoring sensor networks and two in social networks, serves to illustrate the usefulness of the proposed methodology.

Conference Djuric. N., Grbovic M., Vucetic S., Convex Kernelized Sorting, 26th AAAI Conference on Artificial Intelligence (AAAI), Toronto, CA, 2012.

Kernelized sorting is a method for aligning objects across two domains by considering within-domain similarity, without a need to specify a cross-domain similarity measure. In this paper we present the Convex Kernelized Sorting method where, unlike in the previous approaches, the cross-domain object matching is formulated as a convex optimization problem, leading to simpler optimization and global optimum solution. Our method outputs soft alignments between objects, which can be used to rank the best matches for each object, or to visualize the object matching and verify the correct choice of the kernel. It also allows for computing hard one-to-one alignments by solving the resulting Linear Assignment Problem. Experiments on a number of cross-domain matching tasks show the strength of the proposed method, which consistently achieves higher accuracy than the existing methods.

Conference Grbovic, M., Djuric N., Vucetic S., Supervised Clustering of Label Ranking Data, SIAM Conf. on Data Mining (SDM), Anaheim, CA, 2012 (Best of SDM 2012: a top 10 paper).

In this paper we study supervised clustering in the context of label ranking data. Segmentation of such complex data has many potential real-world applications. For example, in target marketing, the goal is to cluster customers in the feature space by taking into consideration the assigned, potentially incomplete product preferences, such that the preferences of instances within a cluster are more similar than the preferences of customers in the other clusters. We establish several heuristic baselines for this application that make use of well-known algorithms such as K-means, and propose a principled algorithm specifically tailored for this type of clustering. It is based on the PlackettLuce (PL) probabilistic ranking model. Each cluster is represented as a union of Voronoi cells defined by a set of prototypes and is assigned a set of PL label scores that determine the cluster-specific label ranking. The unknown cluster PL parameters and prototype positions are determined using a supervised learning technique. Cluster membership and ranking for a new instance is determined by membership of its nearest prototype. The proposed algorithms were empirically evaluated on synthetic and reallife label ranking data. The PL-based method was superior to the heuristically-based supervised clustering approaches. The proposed PL-based algorithm was also evaluated on the task of label ranking prediction. The results showed that it is highly competitive to the state of the art label ranking algorithms, and that it is particularly accurate on data with partial rankings.

Journal Wang, Z., Crammer, K., Vucetic, S., Breaking the Curse of Kernelization: Budgeted Stochastic Gradient Descent for Large-Scale SVM Training, Journal of Machine Learning Research, Vol. 13, pp. 3103−3131, 2012.

Online algorithms that process one example at a time are advantageous when dealing with very large data or with data streams. Stochastic Gradient Descent (SGD) is such an algorithm and it is an attractive choice for online Support Vector Machine (SVM) training due to its simplicity and effectiveness. When equipped with kernel functions, similarly to other SVM learning algorithms, SGD is susceptible to the curse of kernelization that causes unbounded linear growth in model size and update time with data size. This may render SGD inapplicable to large data sets. We address this issue by presenting a class of Budgeted SGD (BSGD) algorithms for large-scale kernel SVM training which have constant space and constant time complexity per update. Specifically, BSGD keeps the number of support vectors bounded during training through several budget maintenance strategies. We treat the budget maintenance as a source of the gradient error, and show that the gap between the BSGD and the optimal SVM solutions depends on the model degradation due to budget maintenance. To minimize the gap, we study greedy budget maintenance methods based on removal, projection, and merging of support vectors. We propose budgeted versions of several popular online SVM algorithms that belong to the SGD family. We further derive BSGD algorithms for multi-class SVM training. Comprehensive empirical results show that BSGD achieves higher accuracy than the state-of-the-art budgeted online algorithms and comparable to non-budget algorithms, while achieving impressive computational efficiency both in time and space during training and prediction.

Journal Grbovic M., Weichang L., Peng X., Usadi A. K., Vucetic S., Decentralized Fault Detection and Diagnosis via Sparse PCA based Decomposition and Maximum Entropy Decision Fusion, Journal of Process Control, Vol. 22, pp. 738– 750, 2012.

This paper proposes an approach for decentralized fault detection and diagnosis in process monitoring sensor networks. The sensor network is decomposed into multiple, potentially overlapping, blocks using the Sparse Principal Component Analysis algorithm. Local predictions are generated at each block using Support Vector Machine classifiers. The local predictions are then fused via a Maximum Entropy algorithm. Empirical studies on the benchmark Tennessee Eastman Process data demonstrated that the proposeddecentralizedapproachachievesaccuracycomparabletothatofthefullycentralizedapproach, while offering benefits in terms of fault tolerance, reusability, and scalability.

Journal Coric, V., Djuric, N., Vucetic, S., Traffic State Estimation from Aggregated Measurements using Signal Reconstruction Techniques, Transportation Research Record: Journal of the Transportation Research Board, Traffic Flow Theory and Characteristics, no. 2315, pp. 121- 130, 2012.

The estimation of the state of traffic provides a detailed picture of the conditions of a traffic network based on limited traffic measurements and,assuch,playsakeyroleinintelligenttransportationsystems.Most oftheexistingstateestimationalgorithmsarebasedonKalmanfiltering and its variants, which, starting from the current estimate, predict the future state and then correct it on the basis of new measurements. Most often, traffic measurements are aggregated over multiple time steps, and this procedure raises the question of how to best use this informationforstateestimation.Astandardapproachthatperformsthecorrection only at the time step when the aggregated measurement is received is suboptimal. Reconstructing the high-resolution measurements from the aggregated ones and using them to correct the state estimates at every time step are proposed. Several reconstruction techniques from signal processing, including kernel regression and a reconstruction approachbasedonconvexoptimization,wereconsidered.Theproposed approach was evaluated on real-world NGSIM data collected at Interstate 101, located in LosAngeles, California. Experimental results show thatsignalreconstructionleadstomoreaccuratetrafficstateestimation as compared with the standard approach for dealing with aggregated measurements.

Journal Wang, Z., Lan, L. Vucetic, S., Mixture Model for Multiple Instance Regression and Applications in Remote Sensing, IEEE Transactions on Geoscience and Remote Sensing, Vol. 50 (6), pp. 2226 – 2237, 2012.

The multiple instance regression (MIR) problem arises when a data set is a collection of bags, where each bag contains multiple instances sharing the identical real-valued label. The goal is to train a regression model that can accurately predict label of an unlabeled bag. Many remote sensing applications can be studied within this setting. We propose a novel probabilistic framework for MIR that represents bag labels with a mixture model. It is based on an assumption that each bag contains a prime instance which is responsible for the bag label. An expectation– maximization algorithm is proposed to maximize the likelihood of the mixture model. The mixture model MIR framework is quite flexible, and several existing MIR algorithms can be described as its special cases. The proposed algorithms were evaluated on synthetic data and remote sensing data for aerosol retrieval and crop yield prediction. The results show that the proposed MIR algorithms achieve higher accuracy than the previous state of the art.

Journal Ristovski, K., Vucetic, S., Obradovic, Z., Uncertainty Analysis of Neural Network-Based Aerosol Retrieval, IEEE Transactions on Geoscience and Remote Sensing, Vol. 50 (2), pp. 409- 414, 2012, 2012.

Neural networks have the ability to represent and learn complex regression functions and are very suitable for retrieval of geophysical parameters from remotely sensed data. Neural networks trained to minimize the mean square error are able to estimate the conditional expectation of target variables. In many remote sensing applications, it is also critical to provide estimates of prediction uncertainty. In this paper, we evaluate an approach that, in addition to training a neural network for retrievals, also trains a neural-network-based estimator of retrieval uncertainty. The uncertainty estimator is built under the assumption that uncertainty is a function of input variables. Themethodologywasevaluatedonaerosol-optical-depthretrieval. The data set consists of 38238 collocated Moderate Resolution Imaging Spectrometer (MODIS) satellite instrument and Aerosol Robotic Network ground-based instrument measurements collected over the entire Earth during two years (in 2005–2006). The results indicate that a neural network ensemble is more accurate than the operational MODIS retrieval algorithm called Collection 5 and that the retrieval uncertainty of the ensemble can be estimated with satisfactory accuracy.

2011

Conference Grbovic, M., Vucetic, S., Tracking Concept Change with Incremental Boosting by Minimization of the Evolving Exponential Loss, European Conf. on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD), Athens, Greece, 2011.

Methodsinvolvingensemblesofclassifiers,suchasbaggingandboosting, are popular due to the strong theoretical guarantees for their performance and their superior results. Ensemble methods are typically designed by assuming the training data set is static and completely available at training time. As such, they are not suitable for online and incremental learning. In this paper we propose IBoost, an extension of AdaBoost for incremental learning via optimization of an exponential cost function which changes over time as the training data changes. The resulting algorithm is flexible and allows a user to customize it based on the computational constraints of the particular application. The new algorithm was evaluated on stream learning in presence of concept change. Experimental results showed that IBoost achieves better performance than the original AdaBoost trained from scratch each time the data set changes, and that it also outperforms previously proposed Online Coordinate Boost, Online Boost and its non-stationary modifications, Fast and Light Boosting, ADWIN Online Bagging and DWM algorithms.

Conference Wang, Z., Djuric, N., Crammer, K. Vucetic, S., Trading Representability for Scalability: Adaptive Multi-Hyperplane Machine for Nonlinear Classification, 17th ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining (KDD), San Diego, CA, 2011.

Support Vector Machines (SVMs) are among the most popular and successful classification algorithms. Kernel SVMs often reach state-of-the-art accuracies, but suffer from the curse of kernelization due to linear model growth with data size on noisy data. Linear SVMs have the ability to efficiently learn from truly large data, but they are applicable to a limited number of domains due to low representational power. To fill the representability and scalability gap between linear and nonlinear SVMs, we propose the Adaptive Multi-hyperplane Machine (AMM) algorithm that accomplishes fast training and prediction and has capability to solve nonlinear classification problems. AMM model consists of a set of hyperplanes (weights), each assigned to one of the multiple classes, and predicts based on the associated class of the weight that provides the largest prediction. The number of weights is automatically determined through an iterative algorithm based on the stochastic gradient descent algorithm which is guaranteed to converge to a local optimum. Since the generalization bound decreases with the number of weights, a weight pruning mechanism is proposed and analyzed. The experiments on several large data sets show that AMM is nearly as fast during training and prediction as the state-of-the-art linear SVM solver and that it can be orders of magnitude faster than kernel SVM. In accuracy, AMM is somewhere between linear and kernel SVMs. For example, on an OCR task with 8 million highly dimensional training examples, AMM trained in 300 seconds on a singlecore processor had 0.54% error rate, which was significantly lower than 2.03% error rate of a linear SVM trained in the same time and comparable to 0.43% error rate of a kernel SVM trained in 2 days on 512 processors. The results indicate that AMM could be an attractive option when solving large-scale classification problems. The software is available at www.dabi.temple.edu/~vucetic/AMM.html.

Conference Malbasa, V., Vucetic., S., Spatially Regularized Logistic Regression for Disease Mapping on Large Moving Population, 17th ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining (KDD), San Diego, CA, 2011.

Spatial analysis of disease risk, or disease mapping, typically relies on information about the residence and health status of individuals from population under study. However, residence information has its limitations because people are exposed to numerous disease risks as they spend time outside of their residences. Thanks to the wide-spread use of mobile phones and GPS-enabled devices, it is becoming possible to obtain a detailed record about the movement of human populations. Availability of movement information opens up an opportunity to improve the accuracy of disease mapping. Starting with an assumption that an individual’s disease risk is a weighted average of risks at the locations which were visited, we show that disease mapping can be accomplished by spatially regularized logistic regression. Due to the inherent sparsity of movement data, the proposed approach can be applied to large populations and over large spatial grids. In our experiments, we were able to map disease for a simulated population with 1.6 million people and a spatial grid with thousand locations in several minutes. The results indicate that movement information can improve the accuracy of disease mapping as compared to residential data only. We also studied a privacy-preserving scenario in which only the aggregate statistics are available about the movement of the overall population, while detailed movement information is available only for individuals with disease. The results indicate that the accuracy of disease mapping remains satisfactory when learning from movement data sanitized in this way.

Journal Djuric, N., Radosavljevic, V., Coric, V., Vucetic, S., Travel Speed Forecasting by Means of Continuous Conditional Random Fields, Transportation Research Record: Journal of the Transportation Research Board, Network Modeling 2011, pp. 131-139, 2011.

This paper explores the application of the recently proposed continuous conditional random fields (CCRF) to travel forecasting. CCRF is a flexible, probabilistic framework that can seamlessly incorporate multiple traffic predictors and exploit spatial and temporal correlations inherently present in traffic data. In addition to improving prediction accuracy, the probabilistic approach provides information about prediction uncertainty. Moreover, information about the relative importance of particular predictor and spatial–temporal correlations can be easily extracted from the model. CCRF is fault-tolerant and can provide predictions even when some observations are missing. Several CCRF models were applied to the problem of travel speed prediction in a range from 10 to 60 min ahead and evaluated on loop detector data from a 5.71-mi section of I-35W in Minneapolis, Minnesota. Several CCRF models, with increasing levels of complexity, are proposed to assess performance of the method better. When these CCRF models were compared with the linear regression model, they reduced the mean absolute error by around 4%. The results imply that modeling spatial and temporal neighborhoods in traffic data and combining various baseline predictors under the CCRF framework can be beneficial.

Journal Lan, L., Vucetic, S., Improving Accuracy of Microarray Classification by a Simple Multi-Task Feature Selection Filter, International Journal of Data Mining and Bioinformatics, Vol. 5 (2), pp. 189-208, 2011.

Leveraging information from the publicly accessible data repositories can be very useful when training a classifier from a small-sample microarray data. To achieve this, we proposed a multi-task feature selection filter that borrows strength from auxiliary microarray data. It uses Kruskal-Wallis test on auxiliary data and ranks genes based on their aggregated p-values. The top-ranked genes are selected as features for the target task classifier. The multi-task filter was evaluated on microarray data related to nine different types of cancers. The results showed that the multi-task feature selection is very successful when applied in conjunction with both single-task and multi-task classifiers.

2010

Conference Wang, Z., Crammer, K., Vucetic, S., “Multi-Class Pegasos on a Budget, Proc. Int. Conf. on Machine Learning (ICML), Haifa, Israel, 2010.

When equipped with kernel functions, online learning algorithms are susceptible to the “curse of kernelization” that causes unbounded growth in the model size. To address this issue, we present a family of budgeted online learning algorithms for multi-class classification which have constant space and time complexity per update. Our approach is based on the multi-class version of the popular Pegasos algorithm. It keeps the number of support vectors bounded during learning through budget maintenance. By treating the budget maintenance as a source of the gradient error, we prove that the gap between the budgeted Pegasos and the optimal solution directly depends on the average model degradation due to budget maintenance. To minimize the model degradation, we study greedy multi-class budget maintenance methods based on removal, projection, and merging of support vectors. Empirical results show that the proposed budgeted online algorithms achieve accuracy comparable to non-budget multi-class kernelized Pegasos while being extremely computationally efficient.

Conference Wang, Z., Vucetic, S., Online Passive-Aggressive Algorithms on a Budget, JMLR W&C Proc. Int. Conf. on Artificial Intelligence and Statistics (AISTATS), Sardinia, Italy, 2010.

In this paper a kernel-based online learning algorithm, which has both constant space and update time, is proposed. The approach is based on the popular online PassiveAggressive (PA) algorithm. When used in conjunction with kernel function, the number of support vectors in PA grows without bounds when learning from noisy data streams. This implies unlimited memory and ever increasing model update and prediction time. To address this issue, the proposed budgeted PA algorithm maintains only a fixed number of support vectors. By introducing an additional constraint to the original PA optimization problem, a closed-form solution was derived for the support vector removal and model update. Using the hinge loss we developed several budgeted PA algorithms that can trade between accuracy and update cost. We also developed the ramp loss versions of both original and budgeted PA and showed that the resulting algorithms can be interpreted as the combination of active learning and hinge loss PA. All proposed algorithms were comprehensively tested on 7 benchmark data sets. The experiments showed that they are superior to the existing budgeted online algorithms. Even with modest budgets, the budgeted PA achieved very competitive accuracies to the non-budgeted PA and kernel perceptron algorithms.

Journal Wang, Z., Vucetic, S., Online Training on a Budget of Support Vector Machines using Twin Prototypes, Statistical Analysis and Data Mining Journal, Vol 3 (3), pp. 149-169, 2010.

This paper proposes twin prototype support vector machine (TVM), a constant space and sublinear time support vector machine (SVM) algorithm for online learning. TVM achieves its favorable scaling by memorizing only a fixed-size data summary in the form of example prototypes and their associated information during training. In addition, TVM guarantees that the optimal SVM solution is maintained on all prototypes at any time. To maximize the accuracy of TVM, prototypes are constructed to approximate the data distribution near the decision boundary. Given a new training example, TVM is updated in three steps. First, the new example is added as a new prototype if it is near the decision boundary. If this happens, to maintain the budget, either the prototype farthest away from the decision boundary is removed or two near prototypes are selected and merged into a single one. Finally, TVM is updated by incremental and decremental techniques to account for the change. Several methods for prototype merging were proposed and experimentally evaluated. TVM algorithms with hinge loss and ramp loss were implemented and thoroughly tested on 12 large datasets. In most cases, the accuracy of low-budget TVMs was comparable with the resource-unconstrained SVMs. Additionally, the TVM accuracy was substantially larger than that of SVM trained on a random sample of the same size. Even larger difference in accuracy was observed when comparing with Forgetron, a popular budgeted kernel perceptron algorithm. As expected, the difference in accuracy between hinge loss and ramp loss TVM was negligible and hinge loss version is preferable due to its lower computational cost. The results illustrate that highly accurate online SVMs could be trained from arbitrary large data streams using devices with severely limited memory budgets.

Journal Radosavljevic, V., Vucetic, S., Obradovic, Z., A Data Mining Technique for Aerosol Retrieval Across Multiple Accuracy Measures, IEEE Geoscience and Remote Sensing Letters, Vol 7 (2), pp. 411 – 415, 2010.

A typical approach in supervised learning is to select an accuracy measure and train a predictor that maximizes it. This can be insufficient in remote-sensing applications where predictor performance is often evaluated over multiple domain-specific accuracy measures. Here, we test the hypothesis that predictors can be trained to maximize performance over multiple accuracy measures. To do this, we evaluate several metalearning algorithms on the problem of aerosol optical depth (AOD) retrieval. The multiple accuracy measures included mean squared error, correlation, relativesquarederror,andfractionofsatisfactorypredictions.The proposed metalearning algorithms have a two-layer architecture, where the first layer consists of multiple neural networks, each trained using a different accuracy measure, and the second layer aggregates decisions of the first layer predictors. To evaluate AOD predictors, we used nearly 70000 collocated data points whose attributeswere radiances,solarandviewangles,andterrainelevation from MODerate resolution Imaging Spectrometer (MODIS) instrument satellite observations and whose target AOD variable was obtained from the ground-based AEROsol robotic NETwork (AERONET) instruments. The data were collected at AERONET locations over the globe in the period between and 2007. AOD prediction accuracies of neural networks were compared to the recently developed operational MODIS C005 retrieval algorithm and to several other data-mining methods. Results showed that neural networks are better at reproducing the test data than the operational retrieval algorithm and that predictors obtained by metalearning are robust over multiple accuracy measures. IndexTerms—Aerosolretrieval,metalearning,neuralnetworks.

2009

Conference Wang, Z., Vucetic, S., Fast Online Training of Ramp Loss Support Vector Machines, IEEE Int’l Conf. on Data Mining (ICDM), Miami, FL, 2009.

A fast online algorithm OnlineSVMR for training Ramp-Loss Support Vector Machines (SVMR s) is proposed. It finds the optimal SVMR for t+1 training examples using SVMR built on t previous examples. The algorithm retains the Karush–Kuhn–Tucker conditions on all previously observed examples. This is achieved by an SMO-style incremental learning and decremental unlearning under the ConcaveConvex Procedure framework. Further speedup of training time could be achieved by dropping the requirement of optimality. A variant, called OnlineASVMR , is a greedy approach that approximately optimizes the SVMR objective function and is suitable for online active learning. The proposed algorithms were comprehensively evaluated on large benchmark data sets. The results demonstrate that OnlineSVMR (1) has the similar computational cost as its offline counterpart; (2) outperforms IDSVM, its competing online algorithm that uses hinge-loss, in terms of accuracy, model sparsity and training time. The experiments on online active learning show that for a fixed number of label queries OnlineASVMR (1) achieves consistently better accuracy than QueryAll and competitive accuracy to Greedy approach; (2) outperforms the active learning version of IDSVM.

Conference Grbovic, M., Vucetic, S., Regression Learning Vector Quantization, IEEE Int’l Conf. on Data Mining (ICDM), Miami, FL, 2009.

Learning Vector Quantization (LVQ) is a popular class of nearest prototype classifiers for multiclass classification. Learning algorithms from this family are widely used because of their intuitively clear learning process and ease of implementation. In this paper we propose an extension of the LVQ algorithm to regression. Just like the LVQ algorithm, the proposed modification uses a supervised learning procedure to learn the best prototype positions, but unlike LVQ algorithm for classification, it also learns the best prototype target values. This results in the effective partition of the feature space, similar to the one the K-means algorithm would make. Experimental results on benchmark datasets showed that the proposed Regression LVQ algorithm performs better than the nearest prototype competitors that choose prototypes randomly or through K-means clustering, classification LVQ on quantized target values, and similarly to the memory-based Parzen Window and Nearest Neighbor algorithms.

Conference Das, D., Obradovic, Z., Vucetic, S., Active Selection of Sensor Sites in Remote Sensing Applications, IEEE Int’l Conf. on Data Mining (ICDM), Miami, FL, 2009.

In a data-mining approach, a model for estimation of Aerosol Optical Depth (AOD) from satellite observations is learned using collocated satellite and groundbased observations. For accurate learning of such a spatiotemporal model, it is important to collect ground-based data from a large number of sites. The objective of this project is to determine appropriate locations for the next set of ground-based data collection sites to maximize accuracy of AOD estimation. Ideally, a new site should capture the most significant unseen aerosol patterns and should be the least correlated with the previously observed patterns. We propose achieving this aim by selecting the locations on which the existing prediction model is the most uncertain. Several criteria were considered for site selection, including uncertainty, spatial diversity, similarity in temporal pattern, and their combination. Extensive experiments on globally distributed data over 90 AERONET sites from the years 2005 and 2006 provide strong evidence that sites selected using the proposed algorithms improve the overall AOD prediction accuracy at a faster rate than those selected randomly or based on spatial diversity among sites.

Conference Wang, Z., Vucetic, S., Twin Vector Machines for Online Learning on a Budget, 2009 SIAM Conf. on Data Mining (SDM), Sparks, NV, 2009.

This paper proposes Twin Vector Machine (TVM), a constant space and sublinear time Support Vector Machine (SVM) algorithm for online learning. TVM achieves its favorable scaling by maintaining only a fixed number of examples, called the twin vectors, and their associated information in memory during training. In addition, TVM guarantees that Kuhn-Tucker conditions are satisfied on all twin vectors at any time. To maximize the accuracy of TVM, twin vectors are adjusted during the training phase to approximate the data distribution near the decision boundary. Given a new training example, TVM is updated in three steps. First, the new example is added as a new twin vector if it is near the decision boundary. If this happens, two twin vectors are selected and merged into a single twin vector to maintain the budget. Finally, TVM is updated by incremental and decremental learning to account for the change. Several methods for twin vector merging were proposed and experimentally evaluated. TVMs were thoroughly tested on 12 large data sets. In most cases, the accuracy of low-budget TVMs was comparable to the state of the art resource-unconstrained SVMs. Additionally, the TVM accuracy was substantially larger than that of SVM trained on a random sample of the same size. Even larger difference in accuracy was observed when comparing to Forgetron, a popular kernel perceptron algorithm on a budget. The results illustrate that highly accurate online SVMs could be trained from large data streams using devices with severely limited memory budgets.

Conference Vucetic, S., Coric, V., Wang, Z., Compressed Kernel Perceptrons, Data Compression Conference (DCC), Snowbird, UT, 2009.

Kernel machines are a popular class of machine learning algorithms that achieve state of the art accuracies on many real-life classification problems. Kernel perceptrons are among the most popular online kernel machines that are known to achieve high-quality classification despite their simplicity. They are represented by a set of B prototype examples, called support vectors, and their associated weights. To obtain a classification, a new example is compared to the support vectors. Both space to store a prediction model and time to provide a single classification scale as O(B). A problem with kernel perceptrons is that the number of support vectors tends to grow without bounds with the number of training examples on noisy data. To reduce the strain at computational resources, budget kernel perceptrons have been developed by upper bounding the number of support vectors. In this work, we proposed a new budget algorithm that upper bounds the number of bits needed to store kernel perceptron. Setting the bitlength constraint could facilitate development of hardware and software implementations of kernel perceptrons on resource-limited devices such as microcontrollers. The proposed compressed kernel perceptron algorithm decides on the optimal tradeoff between number of support vectors and their bit precision. The algorithm was evaluated on several benchmark data sets and the results indicate that it can train highly accurate classifiers even when the available memory budget drops below 1 Kbit. This promising result points to a possibility of implementing powerful learning algorithms even on the most resourceconstrained computational devices.

Conference Grbovic, M., Vucetic, S., Decentralized Estimation Using Learning Vector Quantization, Data Compression Conference (DCC), Snowbird, UT, 2009.Publisher page · DOI 10.1109/dcc.2009.77

Journal Uversky, V.N., Oldfield, C.J., Midic, U., Xie, H., Xue, B., Vucetic, S., Iakoucheva, L.M., Obradovic, Z., Dunker, A.K., Unfoldomics of Human Diseases: Linking Protein Intrinsic Disorder with Diseases, BMC Genomics, 10(Suppl 1):S7, 2009.

BACKGROUND: Intrinsically disordered proteins (IDPs) and intrinsically disordered regions (IDRs) lack stable tertiary and/or secondary structure yet fulfills key biological functions. The recent recognition of IDPs and IDRs is leading to an entire field aimed at their systematic structural characterization and at determination of their mechanisms of action. Bioinformatics studies showed that IDPs and IDRs are highly abundant in different proteomes and carry out mostly regulatory functions related to molecular recognition and signal transduction. These activities complement the functions of structured proteins. IDPs and IDRs were shown to participate in both one-to-many and many-to-one signaling. Alternative splicing and posttranslational modifications are frequently used to tune the IDP functionality. Several individual IDPs were shown to be associated with human diseases, such as cancer, cardiovascular disease, amyloidoses, diabetes, neurodegenerative diseases, and others. This raises questions regarding the involvement of IDPs and IDRs in various diseases. RESULTS: IDPs and IDRs were shown to be highly abundant in proteins associated with various human maladies. As the number of IDPs related to various diseases was found to be very large, the concepts of the disease-related unfoldome and unfoldomics were introduced. Novel bioinformatics tools were proposed to populate and characterize the disease-associated unfoldome. Structural characterization of the members of the disease-related unfoldome requires specialized experimental approaches. IDPs possess a number of unique structural and functional features that determine their broad involvement into the pathogenesis of various diseases. CONCLUSION: Proteins associated with various human diseases are enriched in intrinsic disorder. These disease-associated IDPs and IDRs are real, abundant, diversified, vital, and dynamic. These proteins and regions comprise the disease-related unfoldome, which covers a significant part of the human proteome. Profound association between intrinsic disorder and various human diseases is determined by a set of unique structural and functional characteristics of IDPs and IDRs. Unfoldomics of human diseases utilizes unrivaled bioinformatics and experimental techniques, paves the road for better understanding of human diseases, their pathogenesis and molecular mechanisms, and helps develop new strategies for the analysis of disease-related proteins.

2008

Conference Wang, Z., Radosavljevic, V., Han, B., Obradovic, Z., Vucetic, S., Aerosol Optical Depth Prediction from Satellite Observations by Multiple Instance Regression, Proc. 2008 SIAM Conf. on Data Mining (SDM), Atlanta, GA, 2008.

Aerosols are small airborne particles that both reflect and absorb incoming solar radiation and whose effect on the Earth’s radiation budget is one of the biggest challenges of current climate research. To help address this challenge, numerous satellite sensors are employed to achieve globalscale monitoring of aerosols. Given the satellite measurements, the common objective is prediction of Aerosol Optical Depth (AOD). An important property of AOD is its low spatial variability on a scale of tens of kilometers. On the other hand, satellite sensors gather information in the form of multi-spectral images with high spatial resolution where pixels could be as small as a few hundred meters. Given an accurate ground-based AOD measurement over a specific location and time, all the pixels in the vicinity can be assumed to have the same AOD. If we treat satellite measurement at a single pixel as an instance, all pixels from the neighborhood can be considered as a bag of instances labeled with the same AOD. Given a number of bags obtained at numerous locations and at different times we can treat the problem of AOD prediction from satellite attributes as Multiple Instance Regression (MIR). An important challenge is that because of rapidly changing surface properties attribute values of pixels from a bag can vary a lot. This study evaluated several MIR approaches on several synthetic data sets and on a data set consisting of 800 labeled bags, each containing hundreds of pixel instances observed over the Continental U.S. by the MISR satellite instrument. The results indicate that the most successful MIR approach consists of an iterative procedure that detects and discards outlying instances and trains a predictor on the remaining ones.

Conference Radosavljevic, V., Vucetic, S., Obradovic, Z., Spatio-Temporal Partitioning for Improving Aerosol Prediction Accuracy, Proc. 2008 SIAM Conf. on Data Mining (SDM), Atlanta, GA, 2008.

In supervised learning, on data collected over space and time, different relationships can be found over different spatio-temporal regions. In such situations, an appropriate spatio-temporal data partitioning followed by building specialized predictors could often achieve higher overall prediction accuracy than when learning a single predictor on all the data. In practice, partitions are typically decided based on prior knowledge. As an alternative to domain-based partitioning, we propose a method that automatically discovers a spatio-temporal partitioning through the competition of regression models. The method is evaluated on a challenging problem using satellite observations to predict Aerosol Optical Depth (AOD), which represents the amount of depletion that a beam of radiation undergoes as it passes through the atmosphere. Our experiments used more than 20,000 labeled data points collected during 3 years from more than 100 sites worldwide. Our partitioning-based approach was compared with the recently developed operational AOD prediction algorithm, called C5, which uses domain knowledge for spatio-temporal partitioning of the Earth and implements a region-specific deterministic predictor that utilizes forward simulations from the postulated physical models. The results showed that a neural network predictor trained on all the data has accuracy comparable to C5. When specialized neural network predictors were learned on C5-based partitions, the overall prediction accuracy was not improved. On the other hand, our competition-based spatio-temporal data partitioning approach resulted in large accuracy improvements. The most accurate results were obtained when the data from each site were split into two temporal subsets, one for winter-spring months and another for summer-fall months, and two neural network predictors competed for each identified spatio-temporal subset.

Journal Vucetic, S., Han, B., Mi, W., Li, Z., Obradovic, Z., A Data Mining Approach for the Validation of Aerosol Retrievals, IEEE Geoscience and Remote Sensing Letters, 5 (1), 113-117, 2008.

Operational algorithms for retrieval of aerosols from satellite observations are typically created manually based on the domain knowledge. Validation studies, where the retrievals are compared to the available ground-truth data, are periodically performed with the goal of understanding how to further improve the quality of the retrieval algorithms. This letter describes a data-mining approach aimed to facilitate this highly laborintensive process. It is based on training a neural network for retrieval and comparing its performance with that of the operational algorithm. The situations, where a neural network is more accurate, point to the weaknesses of the operational algorithm that could be corrected. Use of decision trees is proposed to provide easily interpretable descriptions of such situations. The approach was applied on 3646 collocated Moderate Resolution Imaging Spectroradiometer and AERONET observations over the continentalU.S.relatedtotheretrievalofaerosolopticalthickness. The experiments showed that the approach is feasible and that it can be a valuable tool for the domain scientists working on the development of retrieval algorithms.

Journal Krynetskaia, N., Xie, H., Vucetic, S., Obradovic, Z., Krynetskiy, E., High Mobility Group Protein B1 is an Activator of Apoptotic Response to Antimetabolite Drugs, Molecular Pharmacology, 73, 260-269, 2008.

We explored the role of a chromatin-associated nuclear protein high mobility group protein B1 (HMGB1) in apoptotic response to widely used anticancer drugs. A murine fibroblast model system generated from Hmgb1(+)(/)(+) and Hmgb1(-/-) mice was used to assess the role of HMGB1 protein in cellular response to anticancer nucleoside analogs and precursors, which act without destroying the integrity of DNA. Chemosensitivity experiments with 5-fluorouracil, cytosine arabinoside (araC), and mercaptopurine (MP) demonstrated that Hmgb1(-/-) mouse embryonic fibroblasts (MEFs) were 3 to 10 times more resistant to these drugs compared with Hmgb1(+)(/)(+) MEFs. Hmgb1-deficient cells showed compromised cell cycle arrest and reduced caspase activation after treatment with MP and araC. Phosphorylation of p53 at Ser12 (corresponding to Ser9 in human p53) and Ser18 (corresponding to Ser15 in human p53), as well as phosphorylation of H2AX after drug treatment, was reduced in Hmgb1-deficient cells. trans-Activation experiments demonstrated diminished activation of proapoptotic promoters Bax, Puma, and Noxa in Hmgb1-deficient cells after treatment with MP or araC, consistent with reduced transcriptional activity of p53. We have demonstrated for the first time that Hmgb1 is an essential activator of cellular response to genotoxic stress caused by chemotherapeutic agents (thiopurines, cytarabine, and 5-fluorouracil), which acts at early steps of antimetabolite-induced stress by stimulating phosphorylation of two DNA damage markers, p53 and H2AX. This finding makes HMGB1 a potential target for modulating activity of chemotherapeutic antimetabolites. Identification of proteins sensitive to DNA lesions that occur without the loss of DNA integrity provides new insights into the determinants of drug sensitivity in cancer cells.

2007

Journal Xie, H., Vucetic, S., Iakoucheva, L.M., Oldfield, C.J., Dunker, A.K., Uversky, V.N., Obradovic, Z., Functional Anthology of Intrinsic Disorder. I. Biological Processes and Functions of Proteins with Long Disordered Regions, Journal of Proteome Research, 6 (5), 1882 -1898, 2007.

Identifying relationships between function, amino acid sequence, and protein structure represents a major challenge. In this study, we propose a bioinformatics approach that identifies functional

Journal Vucetic, S., Xie, H., Iakoucheva, L.M., Oldfield, C.J., Dunker, A.K., Obradovic, Z., Uversky, V.N., Anthology of Intrinsic Disorder. II. Cellular Components, Domains, Technical Terms, Developmental Processes and Coding Sequence Diversities Correlated with Long Disordered Regions, Journal of Proteome Research, 6 (5), 1899 -1916, 2007.

Biologically active proteins without stable ordered structure (i.e., intrinsically disordered proteins) are attracting increased attention. Functional repertoires of ordered and disordered proteins are very different, and the ability to differentiate whether a given function is associated with intrinsic disorder or with a well-folded protein is crucial for modern protein science. However, there is a large gap between the number of proteins experimentally confirmed to be disordered and their actual number in nature. As a result, studies of functional properties of confirmed disordered proteins, while helpful in revealing the functional diversity of protein disorder, provide only a limited view. To overcome this problem, a bioinformatics approach for comprehensive study of functional roles of protein disorder was proposed in the first paper of this series (Xie, H.; Vucetic, S.; Iakoucheva, L. M.; Oldfield, C. J.; Dunker, A. K.; Obradovic, Z.; Uversky, V. N. Functional anthology of intrinsic disorder. 1. Biological processes and functions of proteins with long disordered regions. J. Proteome Res. 2007, 5, 1882-1898). Applying this novel approach to Swiss-Prot sequences and functional keywords, we found over 238 and 302 keywords to be strongly positively or negatively correlated, respectively, with long intrinsically disordered regions. This paper describes approximately 90 Swiss-Prot keywords attributed to the cellular components, domains, technical terms, developmental processes, and coding sequence diversities possessing strong positive and negative correlation with long disordered regions.

Journal Xie, H., Vucetic, S., Iakoucheva, L.M., Oldfield, C.J., Dunker, A.K., Obradovic, Z., Uversky, V.N., Functional Anthology of Intrinsic Disorder. III. Ligands, Postranslational Modifications and Diseases Associated with Intrinsically Disordered Proteins, Journal of Proteome Research, 6 (5), 1917 -1932, 2007.

Currently, the understanding of the relationships between function, amino acid sequence, and protein structure continues to represent one of the major challenges of the modern protein science. As many as 50% of eukaryotic proteins are likely to contain functionally important long disordered regions. Many proteins are wholly disordered but still possess numerous biologically important functions. However, the number of experimentally confirmed disordered proteins with known biological functions is substantially smaller than their actual number in nature. Therefore, there is a crucial need for novel bionformatics approaches that allow projection of the current knowledge from a few experimentally verified examples to much larger groups of known and potential proteins. The elaboration of a bioinformatics tool for the analysis of functional diversity of intrinsically disordered proteins and application of this data mining tool to >200 000 proteins from the Swiss-Prot database, each annotated with at least one of the 875 functional keywords, was described in the first paper of this series (Xie, H.; Vucetic, S.; Iakoucheva, L. M.; Oldfield, C. J.; Dunker, A. K.; Obradovic, Z.; Uversky, V.N. Functional anthology of intrinsic disorder. 1. Biological processes and functions of proteins with long disordered regions. J. Proteome Res. 2007, 5, 1882-1898). Using this tool, we have found that out of the 710 Swiss-Prot functional

2006

Conference Vucetic, S., A Fast Algorithm for Lossless Compression of Data Tables by Reordering, Data Compression Conference (DCC), Snowbird, Utah, 2006.Publisher page · DOI 10.1109/dcc.2006.1

Journal Han, B., Obradovic, Z., Hu, Z.Z., Wu, C. H., Vucetic, S., Substring Selection for Biomedical Document Classification, Bioinformatics, Vol. 22 (17), pp. 2136-2142, 2006.

MOTIVATION: Attribute selection is a critical step in development of document classification systems. As a standard practice, words are stemmed and the most informative ones are used as attributes in classification. Owing to high complexity of biomedical terminology, general-purpose stemming algorithms are often conservative and could also remove informative stems. This can lead to accuracy reduction, especially when the number of labeled documents is small. To address this issue, we propose an algorithm that omits stemming and, instead, uses the most discriminative substrings as attributes. RESULTS: The approach was tested on five annotated sets of abstracts from iProLINK that report on the experimental evidence about five types of protein post-translational modifications. The experiments showed that Naive Bayes and support vector machine classifiers perform consistently better [with area under the ROC curve (AUC) accuracy in range 0.92-0.97] when using the proposed attribute selection than when using attributes obtained by the Porter stemmer algorithm (AUC in 0.86-0.93 range). The proposed approach is particularly useful when labeled datasets are small.

Journal Han, B., Vucetic, S., Braverman, A., Obradovic, Z., A Statistical Complement to Deterministic Algorithms for the Retrieval of Aerosol Optical Thickness from Radiance Data, Engineering Applications of Artificial Intelligence, Vol. 19, No. 7, pp. 787-795, 2006.

As a complement to the conventional deterministic geophysical algorithms, we consider a faster, but less accurate approach: training regression models to predict aerosol optical thickness (AOT) from radiance data. In our study, neural networks trained on a global data set are employed as a global retrieval method. Inverse distance spatial interpolation and region-specific neural networks trained on restricted, localized areas provide local models. We then develop two integrated statistical methods: local error correction of global retrievals and an optimal weighted average of global and local components. The algorithms are evaluated on the problem of deriving AOT from raw radiances observed by the Multi-angle Imaging SpectroRadiometer (MISR) instrument onboard NASA’s Terra satellite. Integrated statistical approaches were clearly superior to global and local models alone. The best compromise between speed and accuracy was obtained through the weighted averaging of global neural networks and spatial interpolation. The results show that, while much faster, statistical retrievals can be quite comparable in accuracy to the far more computationally demanding deterministic methods. Differences in quality vary with season and model complexity.

Journal Peng, K., Radivojac, P., Vucetic, S., Dunker A.K, Obradovic, Z., Length-Dependent Prediction of Protein Intrinsic Disorder, BMC Bioinformatics, vol. 7 (1), 208, 2006.

BACKGROUND: Due to the functional importance of intrinsically disordered proteins or protein regions, prediction of intrinsic protein disorder from amino acid sequence has become an area of active research as witnessed in the 6th experiment on Critical Assessment of Techniques for Protein Structure Prediction (CASP6). Since the initial work by Romero et al. (Identifying disordered regions in proteins from amino acid sequences, IEEE Int. Conf. Neural Netw., 1997), our group has developed several predictors optimized for long disordered regions (>30 residues) with prediction accuracy exceeding 85%. However, these predictors are less successful on short disordered regions ( 30 residues), respectively. A meta predictor was then trained to integrate the specialized predictors into the final predictor model. As the 10-fold cross-validation results showed, the VSL2 predictors achieved well-balanced prediction accuracies of 81% on both short and long disordered regions. Comparisons over the VSL2 training dataset via 10-fold cross-validation and a blind-test set of unrelated recent PDB chains indicated that VSL2 predictors were significantly more accurate than several existing predictors of intrinsic protein disorder. CONCLUSION: The VSL2 predictors are applicable to disordered regions of any length and can accurately identify the short disordered regions that are often misclassified by our previous disorder predictors. The success of the VSL2 predictors further confirmed the previously observed differences in amino acid compositions and sequence properties between short and long disordered regions, and justified our approaches for modelling short and long disordered regions separately. The VSL2 predictors are freely accessible for non-commercial use at http://www.ist.temple.edu/disprot/predictorVSL2.php.

Journal Radivojac, P., Vucetic, S., O'Connor, T.R., Uversky, V.N., Obradovic, Z., Dunker, A.K., Calmodulin Signaling: Analysis and Prediction of a Disorder-Dependent Molecular Recognition, Proteins: Structure, Function, and Bioinformatics, vol. 63(2), pp. 398-410, 2006.

Calmodulin (CaM) signaling involves important, wide spread eukaryotic protein–protein interactions. The solved structures of CaM associated with several of its binding targets, the distinctive binding mechanism of CaM, and the significant trypsin sensitivity of the binding targets combine to indicate that the process of association likely involves coupled binding and folding for both CaM and its binding targets. Here, we use bioinformatics approaches to test the hypothesis that CaM‐binding targets are intrinsically disordered. We developed a predictor of CaM‐binding regions and estimated its performance. Per residue accuracy of this predictor reached 81%, which, in combination with a high recall/precision balance at the binding region level, suggests high predictability of CaM‐binding partners. An analysis of putative CaM‐binding proteins in yeast and human strongly indicates that their molecular functions are related to those of intrinsically disordered proteins. These findings add to the growing list of examples in which intrinsically disordered protein regions are indicated to provide the basis for cell signaling and regulation.

2005

Conference Vucetic , S., Accuracy-Optimized Quantization for High-Dimensional Data Fusion, Data Compression Conference (DCC), Snowbird, Utah, 2005.Publisher page · DOI 10.1109/dcc.2005.10

Conference Peng, K., Vucetic, S., Obradovic, Z., Correcting Sampling Bias in Structural Genomics through Iterative Selection of Underrepresented Targets, 2005 SIAM Conf. on Data Mining (SDM), Newport Beach, CA, 2005.

In this study we proposed an iterative procedure for correcting sampling bias in labeled datasets for supervised learning applications. Given a much larger and unbiased unlabeled dataset, our approach relies on training contrast classifiers to iteratively select unlabeled examples most highly underrepresented in the labeled dataset. Once labeled, these examples could greatly reduce the sampling bias present in the labeled dataset. Unlike active learning methods, the actual labeling is not necessary in order to determine the most appropriate sampling schedule. The proposed procedure was applied on an important bioinformatics problem of prioritizing protein targets for structural genomics projects. We show that the procedure is capable of identifying protein targets that are underrepresented in current protein structure database, the Protein Data Bank (PDB). We argue that these proteins should be given higher priorities for experimental structural characterization to achieve faster sampling bias reduction in current PDB and make it more representative of the protein space.

Journal Obradovic, Z., Peng, K., Vucetic, S., Radivojac, P., Dunker A.K, Exploiting Heterogeneous Sequence Properties Improves Prediction of Protein Disorder, Proteins: Structure, Function, and Bioinformatics, Vol 61, Suppl 7, pp.176-82, 2005.

During the past few years we have investigated methods to improve predictors of intrinsically disordered regions longer than 30 consecutive residues. Experimental evidence, however, showed that these predictors were less successful on short disordered regions, as observed two years ago during the fifth Critical Assessment of Techniques for Protein Structure Prediction (CASP5). To address this shortcoming, we developed a two-level model called VSL1 (CASP6 id: 193-1). At the first level, VSL1 consists of two specialized predictors, one of which was optimized for long disordered regions (>30 residues) and the other for short disordered regions (< or =30 residues). At the second level, a meta-predictor was built to assign weights for combining the two first-level predictors. As the results of the CASP6 experiment showed, this new predictor has achieved the highest accuracy yet and significantly improved performance on short disordered regions, while maintaining high performance on long disordered regions.

Journal Vucetic, S., Obradovic, Z., Collaborative Filtering Using a Regression-Based Approach, Knowledge and Information Systems, Vol. 7, No. 1, pp. 1-22, 2005.

The task of collaborative filtering is to predict the preferences of an active user for unseen items given preferences of other users. These preferences are typically expressed as numerical ratings. In this paper, we propose a novel regression-based approach that first learns a number of experts describing relationships in ratings between pairs of items. Based on ratings provided by an active user for some of the items, the experts are combined by using statistical methods to predict the user’s preferences for the remaining items. The approach was designed to efficiently address the problem of data sparsity and prediction latency that characterise collaborative filtering. Extensive experiments on Eachmovie and Jester benchmark collaborative filtering data show that the proposed regression-based approach achieves improved accuracy and is orders of magnitude faster than the popular neighbour-based alternative. The difference in accuracy was more evident when the number of ratings provided by an active user was small, as is common for real-life recommendation systems. Additional benefits were observed in predicting items with large rating variability. To provide a more detailed characterisation of the proposed algorithm, additional experiments were performed on synthetic data with second-order statistics similar to that of the Eachmovie data. Strong experimental evidence was obtained that the proposed approach can be applied to data over a large range of sparsity scenarios and is superior to non-personalised predictors even when ratings data are very sparse.

Journal Peng, K., Vucetic, S., Radivojac, P., Brown, CJ, Dunker, AK., Obradovic, Z., Optimizing Long Intrinsic Disorder Predictors with Protein Evolutionary Information, Journal of Bioinformatics and Computational Biology, Vol 3, No. 1, pp. 35-60, 2005.

Protein existing as an ensemble of structures, called intrinsically disordered, has been shown to be responsible for a wide variety of biological functions and to be common in nature. Here we focus on improving sequence-based predictions of long (>30 amino acid residues) regions lacking specific 3-D structure by means of four new neural-network-based Predictors Of Natural Disordered Regions (PONDRs): VL3, VL3H, VL3P, and VL3E. PONDR VL3 used several features from a previously introduced PONDR VL2, but benefitted from optimized predictor models and a slightly larger (152 vs. 145) set of disordered proteins that were cleaned of mislabeling errors found in the smaller set. PONDR VL3H utilized homologues of the disordered proteins in the training stage, while PONDR VL3P used attributes derived from sequence profiles obtained by PSI-BLAST searches. The measure of accuracy was the average between accuracies on disordered and ordered protein regions. By this measure, the 30-fold cross-validation accuracies of VL3, VL3H, and VL3P were, respectively, 83.6 +/- 1.4%, 85.3 +/- 1.4%, and 85.2 +/- 1.5%. By combining VL3H and VL3P, the resulting PONDR VL3E achieved an accuracy of 86.7 +/- 1.4%. This is a significant improvement over our previous PONDRs VLXT (71.6 +/- 1.3%) and VL2 (80.9 +/- 1.4%). The new disorder predictors with the corresponding datasets are freely accessible through the web server at http://www.ist.temple.edu/disprot.

Journal Vucetic, S., Obradovic, Z., Vacic, V., Radivojac, P., Peng, K., Iakoucheva, L.M., Lawson, J.D., Brown, C.J., Sikes, J.G., Newton, C., Dunker, A.K., DisProt: A Database of Protein Disorder, Bioinformatics, Vol. 21, No. 1, 2005.

Summary: The Database of Protein Disorder (DisProt) is a curated database that provides structure and function information about proteins that lack a fixed three-dimensional (3D) structure under putatively native conditions, either in their entirety or in part. Starting from the central premise that intrinsic disorder is an important structural class of protein and in order to meet the increasing interest thereof, DisProt is aimed at becoming a central repository of disorder-related information. For each disordered protein, the database includes the name of the protein, various aliases, accession codes, amino acid sequence, location of the disordered region(s), and methods used for structural (disorder) characterization. If applicable, most entries also list the biological function(s) of each disordered region, how each region of disorder is used for function, as well as provide links to PubMed abstracts and major protein databases.

2004

Conference Radivojac, P., Obradovic, Z., Dunker, AK, Vucetic, S., Characterization of Permutation tests for feature selection, 15th European Conference on Machine Learning (ECML), 2004.

We investigate the problem of supervised feature selection within the filtering framework. In our approach, applicable to the two-class problems, the feature strength is inversely proportional to the p-value of the null hypothesis that its class-conditional densities, p(X|Y = 0) and p(X|Y = 1), are identical. To estimate the p-values, we use Fisher’s permutation test combined with the four simple filtering criteria in the roles of test statistics: sample mean difference, symmetric Kullback-Leibler distance, information gain, and chi-square statistic. The experimental results of our study, performed using naive Bayes classifier and support vector machines, strongly indicate that the permutation test improves the above-mentioned filters and can be used effectively when sample size is relatively small and number of features relatively large.

Journal Radivojac, P., Obradovic, Z., Smith, D.K., Zhu, G., Vucetic, S., Brown, C.J., Lawson, J.D., Dunker, A.K., Protein Flexibility and Intrinsic Disorder, Protein Science, Vol. 13 No. 1, pp. 71- 80, 2004.

Comparisons were made among four categories of protein flexibility: (1) low-B-factor ordered regions, (2) high-B-factor ordered regions, (3) short disordered regions, and (4) long disordered regions. Amino acid compositions of the four categories were found to be significantly different from each other, with high-B-factor ordered and short disordered regions being the most similar pair. The high-B-factor (flexible) ordered regions are characterized by a higher average flexibility index, higher average hydrophilicity, higher average absolute net charge, and higher total charge than disordered regions. The low-B-factor regions are significantly enriched in hydrophobic residues and depleted in the total number of charged residues compared to the other three categories. We examined the predictability of the high-B-factor regions and developed a predictor that discriminates between regions of low and high B-factors. This predictor achieved an accuracy of 70% and a correlation of 0.43 with experimental data, outperforming the 64% accuracy and 0.32 correlation of predictors based solely on flexibility indices. To further clarify the differences between short disordered regions and ordered regions, a predictor of short disordered regions was developed. Its relatively high accuracy of 81% indicates considerable differences between ordered and disordered regions. The distinctive amino acid biases of high-B-factor ordered regions, short disordered regions, and long disordered regions indicate that the sequence determinants for these flexibility categories differ from one another, whereas the significantly-greater-than-chance predictability of these categories from sequence suggest that flexible ordered regions, short disorder, and long disorder are, to a significant degree, encoded at the primary structure level.

2003

Conference Peng, K., Vucetic, S., Han, B., Xie, X., Obradovic, Z., Exploiting Unlabeled Data for Improving Accuracy of Predictive Data Mining, Third IEEE Int'l Conf. on Data Mining (ICDM), Melbourne, FL, 2003.

Predictive data mining typically relies on labeled data without exploiting a much larger amount of available unlabeled data. The goal of this paper is to show that using unlabeled data can be beneficial in a range of important prediction problems and therefore should be an integral part of the learning process. Given an unlabeled dataset representative of the underlying distribution and a K-class labeled sample that might be biased, our approach is to learn K contrast classifiers each trained to discriminate a certain class of labeled data from the unlabeled population. We illustrate that contrast classifiers can be useful in one-class classification, outlier detection, density estimation, and learning from biased data. The advantages of the proposed approach are demonstrated by an extensive evaluation on synthetic data followed by real-life bioinformatics applications for ranking PubMed articles by their relevance to protein disorder and cost-effective enlargement of a disordered protein database.

Conference Vucetic, S., Pokrajac, D., Xie H. and Obradovic, Z., Detection of Underrepresented Biological Sequences Using Class-Conditional Distribution Models, Proc. 2003 SIAM Conf. on Data Mining (SDM), San Francisco, CA, 2003.

A labeled sequence data set related to a certain biological property is often biased and, therefore, does not completely capture its diversity in nature. To reduce this sampling bias problem a data mining procedure is proposed for detecting underrepresented relevant sequences. The procedure is aimed at helping domain experts achieve a cost-effective qualitative enlargement of knowledge through an in-depth study of a small number of statistically underrepresented and functionally interesting sequences. Our procedure consists of: (i) learning a class-conditional distribution model on each class of labeled data; (ii) applying the models to select statistically underrepresented unlabeled sequences; and (iii) automatically evaluating their interestingness. An application of the proposed approach is illustrated on an important problem of increasing the data set of confirmed disordered proteins. The obtained results demonstrate the promise of the proposed approach for an efficient reduction of sampling bias in biological databases.

Journal Obradovic, Z., Peng, K., Vucetic, S., Radivojac, P., Brown C., Dunker A.K, Prediction of Intrinsic Protein Disorder, Proteins: Structure, Function and Genetics, Special Issue on CASP5, Vol. 53, Suppl 6, pp. 566-72, 2003.

Blind predictions of intrinsic order and disorder were made on 42 proteins subsequently revealed to contain 9,044 ordered residues, 284 disordered residues in 26 segments of length residues or less, and 281 disordered residues in disordered segments of length greater than 30 residues. The accuracies of the six predictors used in this experiment ranged from 77% to 91% for the ordered regions and from 56% to 78% for the disordered segments. The average of the order and disorder predictions ranged from 73% to 77%. The prediction of disorder in the shorter segments was poor, from 25% to 66% correct, while the prediction of disorder in the longer segments was better, from 75% to 95% correct. Four of the predictors were composed of ensembles of neural networks. This enabled them to deal more efficiently with the large asymmetry in the training data through diversified sampling from the significantly larger ordered set and achieve better accuracy on ordered and long disordered regions. The exclusive use of long disordered regions for predictor training likely contributed to the disparity of the predictions on long versus short disordered regions, while averaging the output values over 61-residue windows to eliminate short predictions of order or disorder probably contributed to the even greater disparity for three of the predictors. This experiment supports the predictability of intrinsic disorder from amino acid sequence. Proteins 2003;53:566–572.

Journal Vucetic, S., Brown C., Dunker A.K, Obradovic, Z., Flavors of Protein Disorder, Proteins: Structure, Function and Genetics, Vol. 52, pp. 573-584, 2003.

Intrinsically disordered proteins are characterized by long regions lacking 3-D structure in their native states, yet they have been so far associated with 28 distinguishable functions. Previous studies showed that protein predictors trained on disorder from one type of protein often achieve poor accuracy on disorder of proteins of a different type, thus indicating significant differences in sequence properties among disordered proteins. Important biological problems are identifying different types, or flavors, of disorder and examining their relationships with protein function. Innovative use of computational methods is needed in addressing these problems due to relative scarcity of experimental data and background knowledge related to protein disorder. We developed an algorithm that partitions protein disorder into flavors based on competition among increasing numbers of predictors, with prediction accuracy determining both the number of distinct predictors and the partitioning of the individual proteins. Using 145 variously characterized proteins with long (>30 amino acids) disordered regions, 3 flavors, called V, C, and S, were identified by this approach, with the V subset containing 52 segments and 7743 residues, C containing 39 segments and 3402 residues, and S containing 54 segments and 5752 residues. The V, C, and S flavors were distinguishable by amino acid compositions, sequence locations, and biological function. For the sequences in SwissProt and 28 genomes, their protein functions exhibit correlations with the commonness and usage of different disorder flavors, suggesting different flavor-function sets across these protein groups. Overall, the results herein support the flavor-function approach as a useful complement to structural genomics as a means for automatically assigning possible functions to sequences.

2001

Conference Vucetic, S. and Obradovic, Z., Classification on Data with Biased Class Distribution, 12th European Conference on Machine Learning (ECML), pp. 527-538, Freiburg, Germany, 2001.

Labeled data for classification could often be obtained by sampling that restricts or favors choice of certain classes. A classifier trained using such data will be biased, resulting in wrong inference and sub-optimal classification on new data. Given an unlabeled new data set we propose a bootstrap method to estimate its class probabilities by using an estimate of the classifier's accuracy on training data and an estimate of probabilities of classifier's predictions on new data. Then, we propose two methods to improve classification accuracy on new data. The first method can be applied only if a classifier was designed to predict posterior class probabilities where predictions of an existing classifier are adjusted according to the estimated class probabilities of new data. The second method can be applied to an arbitrary classification algorithm, but it requires retraining on the properly resampled data. The proposed bootstrap algorithm was validated through experiments with 500 replicates calculated on 1,000 realizations for each of 16 choices of data set size, number of classes, prior class probabilities and conditional probabilities describing a classifier’s performance. Applications of the proposed methodology to a benchmark data set with various class probabilities on unlabeled data and balanced class probabilities on the training data provided strong evidence that the proposed methodology can be successfully used to significantly improve classification on unlabeled data.

Journal Vucetic, S., Tomsovic, K., Obradovic, Z., Discovering Price-Load Relationships in California’s Electricity Market, IEEE Transactions on Power Systems, Vol. 16, No. 2, pp. 280-286, 2001.

This paper reports on characterizing recent price behavior in the California electricity market. Market participants, that is, producers, consumers and traders, are highly motivated by the potential for profits to develop strategies to explore, and exploit, the limits of system operation. These strategies should be reflected in the market as different price to load relationships. We show that a number of regimes, i.e., characteristic behaviors, exist in the price time series, and provide a brief analysis of each regime. Knowledge of the number of regimes, their characteristics and switching dynamics allows insight into the market and power system performance.

2000

Conference Vucetic, S. and Obradovic, Z., Discovering Homogeneous Regions in Spatial Data through Competition, Proc. 17th Int'l. Conf. on Machine Learning (ICML), pp. 1091-1098, Stanford, CA, 2000.

Here, a supervised machine learning algorithm for the analysis of heterogeneous spatial data is proposed. It is based on partitioning the data set into more homogeneous regions by competition of regression models, linear or nonlinear. The algorithm starts from learning a global model and adds new models into the competition until each model becomes specialized for one of the regions. The competition convergence is proven theoretically. The influence of filtering the competing models' residuals for improving convergence speed and accuracy is also discussed. A number of experiments on artificial and real-life spatial data are performed to validate aspects of the algorithm and illustrate its potential applications. The obtained results provide strong evidence that homogeneous regions can be identified with high accuracy by using the proposed approach even when their observed feature spaces highly overlap.

Journal Vucetic, S., Fiez, T., Obradovic, Z., Examination of the Influence of Data Aggregation and Sampling Density on Spatial Estimation, Water Resources Research, Vol. 36, No. 12, pp. 3721- 3731, 2000. Conference publications

Spatial processes may be sampled by point sampling or by aggregate sampling. If aggregate samples are collected over a regular grid and used to represent the central point of each aggregation area, the aggregate sampling functions as a low‐pass filter and may eliminate aliasing during spatial estimation. To assess potential accuracy improvements, a numerical procedure for calculating the estimation error variance was developed. Analysis of point and block sampling techniques for kriging and inverse distance interpolation showed that for the same sampling density, block sampling provides better estimation. To achieve the same error levels, over 30%–50% more point samples were required than block samples. Furthermore, interpolation of block sampled data resulted in lower error variability and surfaces with more visual appeal.

Work with us

Prospective doctoral students

I enjoy working with students who want to develop new machine-learning methods while working on meaningful real-world problems. A good fit values methodological depth but is also willing to understand the scientific, clinical, engineering, or human context behind a problem.

Strong preparation in computer science, mathematics, statistics, engineering, or another quantitative discipline is useful. I am also open to talking with nontraditional students who can bring a new perspective to our interdisciplinary research projects. Email me your CV and a brief description of your interests.

Research collaborators

Some of our strongest collaborations begin with a domain expert who has an important problem but not yet a neatly defined machine-learning problem: data that are difficult to label, information distributed across several sources, substantial expert knowledge that existing models ignore, or a system whose evaluation requires much more than predictive accuracy.

We equally welcome collaborations with AI and, more broadly, computer science researchers across the whole spectrum of machine-learning topics and AI applications.

Industry and external organizations

We are particularly interested in collaborations where the problem requires research rather than simply the deployment of existing AI tools and pipelines.

Past collaborations have involved topics related to industrial systems, computational advertising, healthcare, defense, communication technology, and other settings where meaningful progress required both domain expertise and methodological innovation. Larger interdisciplinary collaborations, sponsored research, and translational opportunities can also be developed through the Center for Hybrid Intelligence.

About

Biography and background.

I am a Professor of Computer and Information Sciences at Temple University and Director of the Center for Hybrid Intelligence. I joined Temple in 2001 and previously served as Chair and Vice Chair of the department. My research has been supported by NSF, NIH, HHS, other government agencies, and industry, and I regularly teach Temple's graduate Machine Learning course. I also enjoy the bigger questions AI raises, such as the relationship between deep learning and consciousness, which I discussed on a Taste of Science panel in 2021. A few years ago, as a side project for our group, I was very enthusiastic about finding a formula for ranking of CS departments based on citations.

CV (PDF) Google Scholar DBLP ORCID Temple profile Email

Recognition

  • NSF CAREER Award, 2006
  • Member of the team with the highest accuracy in protein disorder prediction, CASP 5, 6, and 7 (2002–2006)
  • Led a team ranked among the top in protein function prediction, CAFA 1 and 2 (2010–2014)
  • Top-ten paper, SIAM International Conference on Data Mining, 2012
  • Outstanding paper award, IJCAI 2013
  • KDD Best Reviewer Award, 2018