Showing posts with label journal papers. Show all posts
Showing posts with label journal papers. Show all posts

Thursday, July 2, 2026

O2O Demand Forecasting @ISF2026, Montreal

My first International Symposium of Forecasting, a.k.a. ISF, was in Rotterdam in 2014, when I made a presentation about Global Energy Forecasting Competitions. I loved the conference and community, so I came back to ISF2015 in Riverside, CA, and the following year in Santander, Spain. Since 2014, I have attended every ISF except the one in Oxford due to COVID. I was elected as a director of International Institute of Forecasters in 2017, and re-elected in 2021. I co-chaired the ISF2025 in Beijing, right after the end of my second (and last) term. 

During past 9 years, I was involved more as an organizer than a conference attendee. My attention was mostly on how well the conference was run, how coherent the program was, and how to better serve the members of IIF. As the premier conference in forecasting, ISF only allows a speaker to make one presentation. I did not really put in too many thoughts about the one and only presentation I can give at ISF. Of course, I chose the easiest route, presenting an energy forecasting talk every time. 

This year, I realized that I now had an opportunity to enjoy ISF as an attendee. Maybe I should try something different for my conference talk. Well, I did. I picked a topic in online take-out food delivery: A three-level hierarchy for O2O demand forecasting. It's based on a paper recently accepted by the International Journal of Forecasting. 

Our IJF paper should be online soon, but I want to share some takeaways outside the paper, which might be beneficial to the forecasters from other fields.

1. It takes some courage to break the status quo.

This IJF paper was related to hierarchical forecasting of online to offline takeout food delivery demand. The original ask was just to develop the forecasts, which could have been done simply by using some off-the-shelf software packages. However, we quickly found out that the existing hierarchy was not optimal for neither business management nor hierarchical forecasting. We then took a detour to investigate the methodologies to construct the hierarchy, which eventually turned into this paper. 

2. The same idea can be applied to seemingly different applications and industries

The core idea of this IJF paper was originally from my master thesis back in 2008. It was later adopted by my wife in her doctoral dissertation in 2010. My master thesis was on electric load forecasting, while her dissertation was on supply chain. The O2O forecasting problem did not even exist at that time! Sometimes we just need to think outside the box, or explore outside the traditional boundaries of the disciplines.


3. It's important to build a comprehensive toolbox early. 

During my years as a student, I took many courses from many different departments. Some were due to degree requirements, while others were not closely related to the degrees I was pursuing. At that time, I often wondered whether those materials were going to be useful. For instance, as of today, I have never directly applied anything I learned w.r.t. assembly language. On the other hand, I have encountered a good number of problems that I could solve easily thanks to the tools I picked up in school. This paper is an example. The methodology we proposed involves clustering. To invent custom solutions to this novel problem, we need some decent understanding of optimization. Without those tools, I would have not even thought about customizing the hierarchy for hierarchical forecasting. 

4. IJF editors and reviewers are awesome!!!

This paper went through many rounds of review. I think it's four or five rounds. If counting the reproducibility check, it's even more. I really appreciate the detailed and constructive criticisms offered by the editors and reviewers. It took us a significant amount of time to revise the paper, which helped improve our work. Unfortunately I don't know who they are due to the double-blind review process. Here I want to express my sincere gratitude for them to invest the time helping us make the paper better.

My talk was very well attended with standing room only. Some folks later told me that they couldn't get into the session because the classroom was so packed. I think I made a good decision to break the status quo, by presenting a non-energy talk at ISF.

Mark your calendar for the 47th ISF, Paphos, Cyprus, June 20th to 23rd, 2027.

Tuesday, January 7, 2020

Forecasting with High Frequency Data: M4 Competition and Beyond

M4 competition was a huge success. The International Journal of Forecasting just published a full issue covering all aspects of the competition. I was honored to be invited by the guest editors to write a commentary paper, which focused on the hourly series of the competition.

According to the organizers (see HERE), the M5 competition is coming soon!

Citation

Tao Hong, "Forecasting with high frequency data: M4 competition and beyond," International Journal of Forecasting, vol.36, no.1, pp.191-194, January, 2020. (ScienceDirect)

Forecasting with High Frequency Data: M4 Competition and Beyond

Tao Hong

Abstract

The M4 competition included 100,000 time series, with the frequencies ranging from yearly to hourly. The team rankings differ notably across frequencies for both point and probabilistic forecasting. I discuss the performances of these methods, with an emphasis on the hourly series of the M4 competition. I also discuss forecasting with high-frequency data in general.

Thursday, October 31, 2019

Descriptive Analytics Based Anomaly Detection for Cybersecure Load Forecasting

Data quality has been a big challenge in load forecasting practice, but an underestimated issue in the academic literature. This work was supported by the U.S. Department of Energy through the Cybersecurity for Energy Delivery Systems Program. We were trying to detect anomalies so that accurate load forecasts can be produced even when the data is contaminated.

Citation

Meng Yue, Tao Hong, and Jianhui Wang, "Descriptive analytics based anomaly detection for cybersecure load forecasting," IEEE Transactions on Smart Grid, vol. 10, no. 6, pp. 5964-5974, November, 2019

Descriptive Analytics Based Anomaly Detection for Cybersecure Load Forecasting

Meng Yue, Tao Hong, and Jianhui Wang

Abstract

As power delivery systems evolve and become increasingly reliant on accurate forecasts, they will be more and more vulnerable to cybersecurity issues. A coordinated data attack by sophisticated adversaries can render existing data corrupt or outlier detection methods ineffective. This would have a very negative impact on operational decisions. The focus of this paper is to develop descriptive analytics-based methods for anomaly detection to protect the load forecasting process against cyberattacks to essential data. We propose an integrated solution (IS) and a hybrid implementation of IS (HIIS) that can detect and mitigate cyberattack induced long sequence anomalies. HIIS is also capable of improving true positive rates and reducing false positive rates significantly comparing with IS. The proposed HIIS can serve as an online cybersecure load forecasting scheme.

Thursday, June 20, 2019

Energy Forecasting in the Big Data World

All papers for the International Journal of Forecasting special section on energy forecasting in the big data world have been published online. Out of 14 papers collected for this special section, eight are from GEFCom2017 documenting winning methods, while the other six non-GEFCom2017 papers cover diverse topics in the areas of energy supply, demand and price forecasting.

The guest editorial is HERE. Below is the list of special section papers:
  1. Tao Hong, Jingrui Xie, and Jonathan Black. Global Energy Forecasting Competition 2017: Hierarchical probabilistic load forecasting
  2. Florian Ziel. Quantile regression for the qualifying match of GEFCom2017 probabilistic load forecasting.
  3. I. Dimoulkas, P. Mazidi, and L. Herre. Neural networks for GEFCom2017 probabilistic load forecasting.
  4. Slawek Smyl and N. Grace Hua. Machine learning methods for GEFCom2017 probabilistic load forecasting.
  5. Andrew J. Landgraf. An ensemble approach to GEFCom2017 probabilistic load forecasting.
  6. Cameron Roach. Reconciled boosted models for GEFCom2017 hierarchical probabilistic load forecasting.
  7. Julian de Hoog and Khalid Abdulla. Data visualization and forecast combination for probabilistic load forecasting in GEFCom2017 final match.
  8. Isao Kanda and J.M. Quintana Veguillas. Data preprocessing and quantile regression for probabilistic load forecasting in the GEFCom2017 final match.
  9. Stephen Haben, Georgios Giasemidis, Florian Ziel, and Siddharth Arora. Short term load forecasting and the effect of temperature at the low voltage level.
  10. Jakob W. Messner and Pierre Pinson. Online adaptive lasso estimation in vector autoregressive models for high dimensional wind power forecasting
  11. Dazhi Yang, Elynn Wu, Jan Kleissl. Operational solar forecasting for the real-time market.
  12. Grzegorz Marcjasz, Bartosz Uniejewski, and Rafał Weron. On the importance of the long-term seasonal component in day-ahead electricity price forecasting with NARX neural networks.
  13. Bartosz Uniejewski, Grzegorz Marcjasz, and Rafał Weron. Understanding intraday electricity markets: Variable selection and very short-term price forecasting using LASSO.
  14. Xuerong Li, Wei Shang, and Shouyang Wang. Text-based crude oil price forecasting: A deep learning approach.

Citation 

Tao Hong and Pierre Pinson, "Energy forecasting in the big data world," International Journal of Forecasting, vol. 35, no. 4, pp. 1387-1388, 2019. 

Energy Forecasting in the Big Data World

Tao Hong and Pierre Pinson

Modern information and communication technologies have brought big data to virtually every segment of the energy and utility industries. While forecasting is an important and necessary step in the data-driven decision-making process, the problem of generating better forecasts in the world of big data is an emerging issue and a challenge to both industry and academia. This special section aims to collect top-quality forecasting articles that document cutting-edge research findings and best practices on a wide range of important business problems in the energy industry. Our emphasis is on big data, such as forecasting with high resolution data, the use of high-dimensional processes, forecasting in real-time, and the use of non-traditional data and variables. 

Monday, April 22, 2019

Combining Weather Stations for Electric Load Forecasting

10 years ago, I started looking into how weather data quality issues affect load forecast accuracy. Later, I found that using data from multiple weather stations can help improve the load forecasts (see this SAS white paper). I also invented a weather station selection methodology to automatically select weather stations for a given load zone. After joining UNC Charlotte, I wrote an IJF paper with two collaborators to introduce that methodology. Nowadays many utilities are using it to select their weather stations. Because that IJF paper is reproducible, I often use it as an entrance exam for prospective students interested in joining BigDEAL.

During the past few years, I have been using that IJF paper as a homework problem in my Energy Analytics class. I have been challenging the students to improve the weather station selection methodology. Although the method is hard to beat, every year some students can turn in something better. Last year, I decided to work with the students in the class to write two papers, one on selecting weather stations, and the other on combining weather stations. Right after I made that decision, Antonio Bracale and Pasquale De Falco invited me to write a paper related to ensemble forecasting for a special issue they were editing. Weather station combination apparently fits the scope very well. Although I believed the research deserves publication with a higher tier journal, I accepted the invitation to make this paper open access, with the hope that those who are using the old methodology can upgrade to this new one with minimal effort.

The peer review process was fairly enjoyable. The paper was submitted on March 18, 2019. The first decision, which was a major revision, was sent back to us on April 1, with comments from three reviewers. Most of the review comments were constructive. None of them were as nonsense as some of the reviewers I encountered at IEEE transactions. We submitted the revision on April 8. The paper was accepted on April 12. The editorial office sent me the edited version for proofread on April 16. I was presently surprised that their copy editor did some wordsmith for us. I submitted the proofread version on April 20. The final version was published on April 21.

Citation

Masoud Sobhani, Allison Campbell, Saurabh Sangamwar, Changlin Li, and Tao Hong, "Combining weather stations for electric load forecasting," Energies, vol. 12, no. 8, pp. 1510, April 2019. (open access)

Combining Weather Stations for Electric Load Forecasting

Masoud Sobhani, Allison Campbell, Saurabh Sangamwar, Changlin Li, and Tao Hong

Abstract

Weather is a key factor affecting electricity demand. Many load forecasting models rely on weather variables. Weather stations provide point measurements of weather conditions in a service area. Since the load is spread geographically, a single weather station may not sufficiently explain the variations of the load over a vast area. Therefore, a proper combination of multiple weather stations plays a vital role in load forecasting. This paper answers the question: given a number of weather stations, how should they be combined for load forecasting? Simple averaging has been a commonly used and effective method in the literature. In this paper, we compared the performance of seven alternative methods with simple averaging as the benchmark using the data of the Global Energy Forecasting Competition 2012. The results demonstrate that some of the methods outperform the benchmark in combining weather stations. In addition, averaging the forecasts from these methods outperforms most individual methods.

Monday, April 8, 2019

Global Energy Forecasting Competition 2017: Hierarchical Probabilistic Load Forecasting

Check out the winning methodologies and data used in GEFCom2017! If you don't have access to ScienceDirect, you can use the dropbox link below to access the data.

Citation

Tao Hong, Jingrui Xie, and Jonathan Black, "Global Energy Forecasting Competition 2017: Hierarchical Probabilistic Load Forecasting," International Journal of Forecasting, in press. (ScienceDirect; Data)


Global Energy Forecasting Competition 2017: Hierarchical Probabilistic Load Forecasting

Tao Hong, Jingrui Xie, and Jonathan Black

Abstract

The Global Energy Forecasting Competition 2017 (GEFCom2017) attracted more than 300 students and professionals from over 30 countries for solving hierarchical probabilistic load forecasting problems. Of the series of global energy forecasting competitions that have been held, GEFCom2017 is the most challenging one to date: the first one to have a qualifying match, the first one to use hierarchical data with more than two levels, the first one to allow the usage of external data sources, the first one to ask for real-time ex-ante forecasts, and the longest one. This paper introduces the qualifying and final matches of GEFCom2017, summarizes the top-ranked methods, publishes the data used in the competition, and presents several reflections on the competition series and a vision for future energy forecasting competitions.

Monday, February 11, 2019

Short-term Industrial Reactive Power Forecasting

Two years ago, I started collaborating with a team of Italian researchers. We had our first joint paper on short-term industrial load forecasting published at the 2017 ISGT-Europe. The complete story is HERE.

Since then, we've continued our collaboration. In this paper, we used the data from the same Italian factory. Now we focus on reactive power forecasting, a rarely touched topic in the load forecasting literature. 

Citation

Antonio Bracale, Guido Carpinelli, Pasquale De Falco, and Tao Hong, "Short-Term Industrial Reactive Power Forecasting," International Journal of Electrical Power & Energy Systems, vol.107, pp 177-185, May 2019 (ScienceDirect)

Short-term Industrial Reactive Power Forecasting

Antonio Bracale, Guido Carpinelli, Pasquale De Falco, and Tao Hong

Abstract

Reactive power forecasting is essential for managing energy systems of factories and industrial plants. However, the scientific community has devoted scant attention to industrial load forecasting, and even less to reactive power forecasting. Many challenges in developing a short-term reactive power forecasting system for factories have rarely been studied. Industrial loads may depend on many factors, such as scheduled processes and work shifts, which are uncommon or unnecessary in classical load forecasting models. Moreover, the features of reactive power are significantly different from active power, so some commonly used variables in classical load forecasting models may become meaningless for forecasting reactive power. In this paper, we develop several models to forecast industrial reactive power. These models are constructed based on two forecasting techniques (e.g., multiple linear regression and support vector regression) and two variable selection methods (e.g., cross validation and least absolute shrinkage and selection operator). In the numerical applications based on real data collected from an Italian factory at both aggregate and individual load levels, the proposed models outperform four benchmark models in short forecast horizons.

Tuesday, October 16, 2018

Robust Regression Models for Load Forecasting

One of my doctoral majors is operations research, for which I took many courses in graduate school to build my knowledge in optimization. The topic of my dissertation was on load forecasting. Only two chapters were related to optimization, one on Artificial Neural Networks, and the other on Fuzzy Regression (or Possibilistic Linear Regression).

In fact, the fuzzy regression chapter was the only one that seriously required some optimization skills, which was published as an FODM paper three years after my graduation. To build a fuzzy regression model, I had to formulate the parameter estimation process as a linear program, and solve it in CPLEX. At that time Gurobi was not even able to provide a feasible solution for my fuzzy regression model with 200+ parameters.

After that, I continued my profession in forecasting. I knew my optimization background is helpful to forecasting, but I didn't really expect to apply many optimization skills in forecasting.

About a year ago, we performed a benchmark study to show that four representative load forecasting models would fail miserably with bad input data. That study was published as an IJF paper early this year. At the end of that IJF paper, we mentioned a future research direction of designing more robust load forecasting models.

In this paper, we propose three robust regression models for load forecasting. While all of them are more robust than the ones compared in the IJF paper, the L1 regression model outperform the others. In fact L1 regression is not really new to load forecasting. It has been used for forecast combination, where some people call it Least Absolute Deviation (LAD) regression. Its "general" form, quantile regression, is heavily used in probabilistic load forecasting.
What's new about the L1 regression model in this paper?
We built an L1 regression model with hundreds of parameters. In fact it shares the same variable combination as the Vanilla model used in Global Energy Forecasting Competitions. Building such a model is nontrivial. We didn't find an off-the-shelf package to do what we need, so we formulated it as a linear program and solved it using MATLAB's linprog.
Among hundreds of techniques that are applicable to load forecasting, how did I find L1 regression?
The idea didn't come from nowhere. When I was working on my doctoral dissertation at FANGroup (Fuzzy And Neural Group), a few other students were working on another project sponsored by U.S. Army Research Office. They were investigating some features and applications of l1 norm. Although I was thinking about applying l1 norm to load forecasting, I didn't find a good use case at that time.

Well, it's better late than never. The skills I acquired 10 years ago came handy for this paper.

Citation

Jian Luo, Tao Hong, and Shu-Cherng Fang, "Robust regression models for load forecasting," submitted to IEEE Transactions on Smart Grid, in press.


Robust Regression Models for Load Forecasting

Jian Luo, Tao Hong, and Shu-Cherng Fang

Abstract

Electric load forecasting has been extensively studied during the past century. While many models and their variants have been proposed and tested in the load forecasting literature, most of the existing case studies have been conducted using the data collected under normal operating conditions. A recent case study shows that four representative load forecasting models easily fail under data integrity attacks. To address this challenge, we propose three robust load forecasting models including two variants of the iteratively re-weighted least squares regression models and an L1 regression model. Numerical experiments indicate the dominating performance of the three proposed robust regression models, especially L1 regression, compared to other representative load forecasting models. 

Monday, July 9, 2018

From Club Convergence of Per Capita Industrial Pollutant Emissions to Industrial Transfer Effects: An Empirical Study Across 285 Cities in China

China has grown to the world's second largest economy by nominal GDP. Many factors attribute to such rapid growth, such as globalization and hard-working Chinese people. Nevertheless, we can't ignore the pollution resulted from the industrialization. Dr. Chang Liu brought the research problem to me when she visited BigDEAL last year. We spent a year investigating the relationship between industrial transfer effects and per capita industrial pollutant emissions across 285 cities in China. We identified four convergence clubs for SO2 emissions, and three convergence clubs for soot emissions. We also concluded that industrial transfer effects can lead to multiple steady-state equilibria. This presents some evidence to support region-specific environmental policies and execution strategies. 

This is the first time I sent a paper to Energy Policy. The original version was submitted on Feb 5, 2018. Within five months, the paper was published after three revisions. The entire publication process was quite pleasant.

Citation
Chang Liu, Tao Hong, Huaifeng Liu, and Lili Wang, "From club convergence of per capita industrial pollutant emissions to industrial transfer effects: an empirical study across 285 cities in China," Energy Policy, vol.121, pp 300-313, October 2018. (ScienceDirect)

From Club Convergence of Per Capita Industrial Pollutant Emissions to Industrial Transfer Effects: An Empirical Study Across 285 Cities in China

Chang Liu, Tao Hong, Huaifeng Liu, and Lili Wang

Abstract

The process of industrialization has led to an increase in air pollutant emissions in China. At the regional level, industrial restructuring and industrial transfer from eastern China to western China have caused a significant difference in pollutant emissions among various cities. This paper analyzes per capita industrial pollutant emissions across 285 prefecture-level cities from 2003 to 2015, aiming to reveal how industrial transfer affects the formation of convergence clubs. Whether industrial pollutant emissions across heterogeneous cities converge to a unique steady-state equilibrium is first identified based on the concept of club convergence. Logit regression analysis is then applied to assess the effects of industrial transfer on the observed clubs. The log t-test highlights four convergence clubs for industrial SO2 emissions and three clubs for industrial soot emissions. The regression analysis results reveal that the effects of industrial transfer can lead to multiple steady-state equilibria, suggesting region-specific environmental policies and execution strategies. In addition, accelerating the development of clean energy technologies in emission-intense regions should be further emphasized. 

Monday, June 18, 2018

Combining Probabilistic Load Forecasts

We often find simple averaging as a plausible solution for combining point forecasts. Combining probabilistic forecasts is not that trivial. The literature of combining probabilistic load forecasts is rather limited. Previously, we developed a Quantile Regression Averaging (QRA) method to generate probabilistic load forecasts by combining point forecasts. This work is a follow up, where we combine probabilistic load forecasts to generate a more accurate probabilistic forecast. The method we proposed here is a Constrained Quantile Regression Averaging (CQRA) method, where the parameters of a quantile regression model are non-negative and sum up to 1. We applied the method to loads at both high voltage level and household level, showing better results than the benchmarks.

Among my papers published so far, this one has the shortest title.

Citation
Yi Wang, Ning Zhang, Yushi Tan, Tao Hong, Daniel Kirschen, and Chongqing Kang, "Combining probabilistic load forecasts," IEEE Transactions on Smart Grid, in press, available online. (arXiv; IEEE Xplore).

Combining Probabilistic Load Forecasts

Yi Wang, Ning Zhang, Yushi Tan, Tao Hong, Daniel Kirschen, and Chongqing Kang

Abstract

Probabilistic load forecasts provide comprehensive information about future load uncertainties. In recent years, many methodologies and techniques have been proposed for probabilistic load forecasting. Forecast combination, a widely recognized best practice in point forecasting literature, has never been formally adopted to combine probabilistic load forecasts. This paper proposes a constrained quantile regression averaging (CQRA) method to create an improved ensemble from several individual probabilistic forecasts. We formulate the CQRA parameter estimation problem as a linear program with the objective of minimizing the pinball loss and the constraints that the parameters are nonnegative and summing up to one. We demonstrate the effectiveness of the proposed method using two publicly available datasets, the ISO New England data and Irish smart meter data. Comparing with the best individual probabilistic forecast, the ensemble can reduce the pinball score by 4.39% on average. The proposed ensemble also demonstrates superior performance over nine other benchmark ensembles.

Thursday, June 14, 2018

A Semi-heterogeneous Approach to Combining Crude Oil Price Forecasts

Forecast combination is an effective method to enhance the accuracy. Most combination methods in the literature can be grouped two categories, heterogeneous combination and homogeneous combination, with each having pros and cons. I collaborated with my former visiting scholar Dr. Jue Wang and her colleagues to develop a semi-heterogeneous approach to combining forecasts. We leveraged the decomposition-reconstruction concept, mixing and matching 4 decomposition methods with 4 forecasting techniques. In total this process generates 16 forecasts for combination, which is easier than applying 16 completely different techniques (a.k.a. heterogeneous combination) and more robust than producing 16 different forecasts from one technique (a.k.a. homogeneous combination). Furthermore, the proposed method leads to more accurate forecasts than its counterparts.

Citation
Jue Wang, Xiang Li, Tao Hong, and Shouyang Wang, "A semi-heterogeneous approach to combining crude oil price forecasts," Information Sciences, vol.460-461, pp 279-292, September 2018. (ScienceDirect)


A Semi-heterogeneous Approach to Combining Crude Oil Price Forecasts

Jue Wang, Xiang Li, Tao Hong, and Shouyang Wang

Abstract

Crude oil price forecasting has received increased attentions due to its significant role in the global economy. Accurate crude oil price forecasts often lead to a rapid new production development with higher quality and less cost. Making such accurate forecasts, however, is challenging due to the intrinsic complexity of oil market mechanism. Many techniques have been tested in the crude oil price forecasting literature. Although forecast combination is a well-known method to improve forecast accuracy, generating forecasts using various techniques tend to be labor intensive. How to efficiently generate many individual forecasts for combination becomes a research question in crude oil price forecasting. Recently, several signal decomposition methods have been suggested for processing the oil price signals. In this paper, we propose a semi-heterogeneous approach to combining crude oil price forecasts, which interacts a set of decomposition methods with a set of forecasting techniques. We first decompose the original price series using four decomposition methods, such as Wavelet Analysis, Singular Spectral Analysis, Empirical Mode Decomposition, and Variational Mode Decomposition. We then use four different forecasting techniques, such as Autoregressive Models, Autoregressive Integrated Moving Average Models, Artificial Neural Networks, and Support Vector Regression Models, to forecast the components from each decomposition methods. Finally, we reconstruct the price forecasts from the forecasted components. This process generates 16 price forecasts in total for combination. We test the combination based on all individual forecasts, as well as a subset of the individual forecasts selected using Tabu Search. The experimental results demonstrate that the forecasting models with the addition of a decomposition technique can have an error reduction of 30.6% compared to benchmark models on average. The combined forecasts outperform the individual forecasts on average. Furthermore, comparing with the heterogeneous combination of 4 individual forecasts, the semi-heterogeneous combinations reduce the errors by 56.6% (w/o Tabu Search) and 61.6% (w/ Tabu Search).

Friday, March 23, 2018

Review of Smart Meter Data Analytics: Applications, Methodologies, and Challenges

About 10 years ago, the term "smart grid" was officially defined in the Energy Independence and Security Act of 2007 (EISA-2007). Soon after that, many power companies started their smart meter deployments. As of 2016, more than 70 million smart meters were installed in the united states. The world installed base was projected to reach 780 million by 2020, pushed by the mass roll-outs in China. We are now sitting on a gold mine of data collected by these smart meters. Last year I gave a forecast:
The energy companies will be moving more Gigabytes of data than GWh of electricity.
The scientific community has been trying to understand the smart meter data and get some actionable insights out of it. Thousands of papers have been published in the recent decade on the various aspects of smart meter data analytics. Last year, I worked with my collaborators in Tsinghua University to complete a review of smart meter data analytics. The paper was just put on the IEEE Xplore yesterday.

This is the longest paper ever published by the IEEE Transactions on Smart Grid. I'm sure reading this 24-page review article can save the readers significant amount of time from digging thousands of papers in the literature. Load forecasters may find some interesting stuff in Section III, which is dedicated to load forecasting in the smart grid era.

Citation

Yi Wang, Qixin Chen, Tao Hong, and Chongqing Kang, "Review of smart meter data analytics: applications, methodologies, and challenges," IEEE Transactions on Smart Grid, in press. (working paper; IEEE Xplore)

Review of Smart Meter Data Analytics: Applications, Methodologies, and Challenges

Yi Wang, Qixin Chen, Tao Hong, and Chongqing Kang

Abstract

The widespread popularity of smart meters enables an immense amount of fine-grained electricity consumption data to be collected. Meanwhile, the deregulation of the power industry, particularly on the delivery side, has continuously been moving forward worldwide. How to employ massive smart meter data to promote and enhance the efficiency and sustainability of the power grid is a pressing issue. To date, substantial works have been conducted on smart meter data analytics. To provide a comprehensive overview of the current research and to identify challenges for future research, this paper conducts an application-oriented review of smart meter data analytics. Following the three stages of analytics, namely, descriptive, predictive and prescriptive analytics, we identify the key application areas as load analysis, load forecasting, and load management. We also review the techniques and methodologies adopted or developed to address each application. In addition, we also discuss some research trends, such as big data issues, novel machine learning technologies, new business models, the transition of energy systems, and data privacy and security.

Thursday, February 8, 2018

Real-time Anomaly Detection for Very Short-term Load Forecasting

Many very short-term load forecasting (VSTLF) models in literature rely on lagged loads, while most of these VSTLF papers assume perfect information of the lagged loads. As a result, the accuracy reported in the VSTLF literature has been amazingly high. In reality, however, load forecasters may not have access to the load values of the most recent few hours. The imperfection of the recent load information would certainly affect the load forecast accuracy. This paper tackles a practical problem, how to detect the anomalies in the most recent load information.

Citation
Jian Luo, Tao Hong and Meng Yue, "Real-time anomaly detection for very short-term load forecasting," Journal of Modern Power Systems and Clean Energy, in press, available online. (open access)

Real-time Anomaly Detection for Very Short-term Load Forecasting

Jian Luo, Tao Hong and Meng Yue

Abstract

Although the recent load information is critical to very short-term load forecasting (VSTLF), power companies often have difficulties in collecting the most recent load values accurately and timely for VSTLF applications. This paper tackles the problem of real-time anomaly detection in most recent load information used by VSTLF. This paper proposes a model-based anomaly detection method that consists of two components, a dynamic regression model and an adaptive anomaly threshold. The case study is developed using the data from ISO New England. This paper demonstrates that the proposed method significantly outperforms three other anomaly detection methods including two methods commonly used in the field and one state-of-the-art method used by a winning team of the Global Energy Forecasting Competition 2014. Finally, a general anomaly detection framework is proposed for the future research. 

Friday, February 2, 2018

Load Forecasting Using 24 Solar Terms

Can we use Chinese calendar to forecast the load in the U.S.? Since I started my load forecasting practice 10 years ago, this has been a question sitting in my mind. One year ago, we decided to the test this idea. In short, the answer is YES. In the big data era, this approach would fall in the category of leveraging a variety of data sources.

This paper will be collected in the MPCE special section "Forecasting in Modern Power Systems" (Call For Papers). The paper is open access, so you can read the full content and download the PDF file for free. Special thanks to the journal editorial office for the neat copy-editing work. I truly enjoyed the publication process. Unlike most other open access journals that charge the authors a big fee for publishing the papers, this one does not charge a dime. I would highly recommend this journal to those who are interested in publishing open access papers in energy forecasting but do not want to pay for the publication fees. 

Citation
Jingrui Xie and Tao Hong, "Load forecasting using 24 solar terms," Journal of Modern Power Systems and Clean Energy, in press, available online. (open access

Load Forecasting Using 24 Solar Terms

Jingrui Xie and Tao Hong

Abstract

Calendar is an important driving factor of electricity demand. Therefore, many load forecasting models would incorporate calendar information. Frequently used calendar variables include hours of a day, days of a week, months of a year, and so forth. During the past several decades, a widely-used calendar in load forecasting is the Gregorian calendar from the ancient Rome, which dissects a year into 12 months based on the Moon’s orbit around the Earth. The applications of alternative calendars have rarely been reported in the load forecasting literature. This paper aims at discovering better means than Gregorian calendar to categorize days of a year for load forecasting. One alternative is the solar-term calendar, which divides the days of a year into 24 terms based on the Sun’s position in the zodiac. It was originally from the ancient China to guide people for their agriculture activities. This paper proposes a novel method to model the seasonal change for load forecasting by incorporating the 24 solar terms in regression analysis. The case study is conducted for the eight load zones and the system total of ISO New England. Results from both cross-validation and sliding simulation show that the forecast based on the 24 solar terms is more accurate than its counterpart based on the Gregorian calendar.

Tuesday, January 30, 2018

Integrated Facility Location and Production Scheduling in Multi-generation Energy Systems

Multiple energy systems integration has been attracting much attention during the recent years. At the PES General Meeting 2017, I co-chaired a session "Accommodating Intermittent Renewable Energy by Multiple Energy Systems Integration: Forecasting, Operations and Planning", which was very well attended.

In this paper, we tackle the problem from a systems perspective by investigating the network design philosophy and exploring the economic value of the multi-generation technologies under demand uncertainties. The work is certainly out of the mainstream energy forecasting research I have published in the past. Nevertheless, readers of this blog may find it intriguing, because some of the forecasts we produce are going to be fed into the decision making processes for locating power plants and generation scheduling.

Citation

Qiaochu He and Tao Hong, "Integrated facility location and production scheduling in multi-generation energy systems," Operations Research Letters, vol.46, no.1, pp 153-157, January 2018. (working paper; ScienceDirect)

Integrated Facility Location and Production Scheduling in Multi-generation Energy Systems

Qiaochu He and Tao Hong

Abstract

In this paper, we investigate the energy system design problems with the multi-generation technologies, i.e., simultaneous generation of multiple types of energy. Our results illustrate the economic value of multi-generation technologies to reduce spatio-temporal demand uncertainty by risk pooling both within and across different facilities.

Wednesday, August 9, 2017

Benchmarking Robustness of Load Forecasting Models under Data Integrity Attacks

GIGO - garbage in, garbage out. In forecasting, GIGO means that if the model is fed with garbage (bad) data, the forecast would be bad too. In the power industry, bad load forecasts often result in waste of energy resources, financial losses, brownouts or even blackouts.

Anomaly detection and data cleansing procedures may help alleviate some of the bad data from the input side. However, what if the bad data was created by hackers? Can the existing models "survive" or stay accurate under data attacks? This paper offers some benchmark results.

This paper sets a few "first":
  1. This is the first paper formally addressing the cybersecurity issues in the load forecasting literature. I believe that the data attacks should be of a great concern to the forecasting community. This paper sets a solid ground for future research. 
  2. This is my first journal paper co-authored with a professor in my doctoral committee, Dr. Shu-Cherng Fang. Many years ago, I picked up the topic of my dissertation from one of my consulting projects. I then invited a team of world class professors from different areas to form the committee. None of them were really into load forecasting, though I had the opportunities learning from different perspectives. 
  3. This is my first journal paper that went through one year of peer review cycle, the longest peer review I've experienced. It's definitely worth the effort. The IJF editors and reviewers certainly spent a significant amount of time reading the paper and offered so many constructive comments. I wish I could know their names and identify them in the acknowledgement section.
Citation

Jian Luo, Tao Hong and Shu-Cherng Fang, "Benchmarking robustness of load forecasting models under data integrity attacks", International Journal of Forecasting, accepted. (working paper)


Benchmarking robustness of load forecasting models under data integrity attacks

Jian Luo, Tao Hong and Shu-Cherng Fang

Abstract

As the internet continues to expand its footprint, cybersecurity has become a major concern for the governments and private sectors. One of the cybersecurity issues is on data integrity attacks. In this paper, we focus on the power industry, where the forecasting processes heavily rely on the quality of data. The data integrity attacks are expected to harm the performance of forecasting systems, which greatly impact the financial bottom line of power companies and the resilience of power grids. Here we reveal how data integrity attacks can affect the accuracy of four representative load forecasting models (i.e., multiple linear regression, support vector regression, artificial neural networks, and fuzzy interaction regression). We first simulate some data integrity attacks by randomly injecting some multipliers that follow a normal or uniform distribution to the load series. Then the aforementioned four load forecasting models are applied to generate one-year ahead ex post point forecasts for comparisons of their forecast errors. The results show that the support vector regression model, trailed closely by the multiple linear regression model, is most robust, while the fuzzy interaction regression model is least robust among the four. Nevertheless, all of the four models fail to provide satisfying forecasts when the scale of data integrity attacks becomes large. This presents a serious challenge to the load forecasters and the broader forecasting community: How to generate accurate forecasts under data integrity attacks? We use the publicly-available data from Global Energy Forecasting Competition 2012 to construct the case study. At the end, we also offer an outlook of potential research topics for future studies.

Tuesday, August 1, 2017

Variable Selection Methods for Probabilistic Load Forecasting: Empirical Evidence from Seven states of the United States

I am an evidence-based man. This mentality saves me tremendous amount of time in recent years. I have been minimizing my time in following bluffs in the literature. On the other hand, I have been developing empirical case studies and encourage the community to contribute to he empirical research.

In my GEFCom2014 paper, I raised the following question to the forecasting community:
Can a better point forecasting model lead to a better probabilistic forecast?
To answer this question, we have to first understand the definition of "better", a.k.a., forecast evaluation measures and methods. In this paper, we compared two variable selection methods based on point and probabilistic error measures respectively. The case study covers seven states of the US. The results from this paper can hopefully be leveraged by future empirical studies for comparison purposes.

Wednesday, May 17, 2017

Wind Speed for Load Forecasting Models

One way to categorize the load forecasting papers is based on the variables used in those forecasting models. Because many people who wrote load forecasting papers only had access to the load data with time stamps, they had to propose the models based on the load series only. The representative techniques include exponential smoothing and the ARIMA family. Sometimes people also include the calendar information to come up with some regression models with classification variables. Although these are good and powerful techniques, their real-world applications in load forecasting are very limited. I have criticized those "load-only" models in some of my papers, such as the IJF2016 paper on recency effect:
Both seasonal naïve models perform very poorly compared with the other four models. Seasonal naïve models are used commonly for benchmarking purposes in other industries, such as the retail and manufacturing industries. In load forecasting, the two applications in which seasonal naïve models are most useful are: (1) benchmarking the forecast accuracy for very unpredictable loads, such as household level loads; and (2) comparisons with univariate models. In most other applications, however, the seasonal naïve models and other similar naïve models are not very meaningful, due to the lack of accuracy. 

Wednesday, August 24, 2016

Guest Editorial: Big Data Analytics for Grid Modernization

IEEE Transactions on Smart Grid just published our special section on Big Data Analytics for Grid Modernization. The guest editorial is on IEEE Xplore with open access. The original Call for Papers is HERE.

While "big data" is quickly becoming a buzz word (see THIS POST), in this guest editorial we discussed our interpretation from four aspects:
  1. The data involved in the analysis is big in at least one of its three defining dimensions: volume, variety or velocity. The big data used in the utility industry includes but is not limited to smart meter data, phasor measurement unit data, weather data, and social media data. 
  2. The problem under investigation is to prepare for analyzing the big data, such as data compression and data security issues. 
  3. The methodology requires customized modeling of individual components of a system, or leads to in-depth understanding of the individual components. For instance, estimating the invisible solar generation belongs to this category. 
  4. The technology can be used to help reach the answer faster, or answer the questions otherwise difficult to answer. For example, a distributed platform can be used to speed up the analytic tasks.
Thanks to the diligent work from our guest editors, reviewers and the authors, we are able to present a high-quality collection of papers to the community. Below is the list of 17 special section papers:
  1. D. Zhou, J. Guo, Y. Zhang, J. Chai, H. Liu, Y. Liu, C. Huang, X. Gui and Y. Liu, "Distributed Data Analytics Platform for Wide-Area Synchrophasor Measurement Systems," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2397-2405, Sept. 2016
  2. P. H. Gadde, M. Biswal, S. Brahma and H. Cao, "Efficient Compression of PMU Data in WAMS," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2406-2413, Sept. 2016
  3. X. Tong, C. Kang and Q. Xia, "Smart Metering Load Data Compression Based on Load Feature Identification," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2414-2422, Sept. 2016
  4. J. Hu and A. V. Vasilakos, "Energy Big Data Analytics and Security: Challenges and Opportunities," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2423-2436, Sept. 2016
  5. Y. Wang, Q. Chen, C. Kang and Q. Xia, "Clustering of Electricity Consumption Behavior Dynamics Toward Big Data Applications," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2437-2447, Sept. 2016
  6. S. Ben Taieb, R. Huser, R. J. Hyndman and M. G. Genton, "Forecasting Uncertainty in Electricity Smart Meter Data by Boosting Additive Quantile Regression," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2448-2455, Sept. 2016
  7. H. Shaker, H. Zareipour and D. Wood, "Estimating Power Generation of Invisible Solar Sites Using Publicly Available Data," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2456-2465, Sept. 2016
  8. H. Shaker, H. Zareipour and D. Wood, "A Data-Driven Approach for Estimating the Power Generation of Invisible Solar Sites," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2466-2476, Sept. 2016
  9. X. Zhang and S. Grijalva, "A Data-Driven Approach for Detection and Estimation of Residential PV Installations," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2477-2485, Sept. 2016
  10. H. Wang and J. Huang, "Cooperative Planning of Renewable Generations for Interconnected Microgrids," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2486-2496, Sept. 2016
  11. J. Peppanen, M. J. Reno, R. J. Broderick and S. Grijalva, "Distribution System Model Calibration With Big Data From AMI and PV Inverters," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2497-2506, Sept. 2016
  12. Y. C. Chen, J. Wang, A. D. Domínguez-García and P. W. Sauer, "Measurement-Based Estimation of the Power Flow Jacobian Matrix," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2507-2515, Sept. 2016
  13. H. Sun, Z. Wang, J. Wang, Z. Huang, N. Carrington and J. Liao, "Data-Driven Power Outage Detection by Social Sensors," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2516-2524, Sept. 2016
  14. H. Jiang, X. Dai, D. W. Gao, J. J. Zhang, Y. Zhang and E. Muljadi, "Spatial-Temporal Synchrophasor Data Characterization and Analytics in Smart Grid Fault Detection, Identification, and Impact Causal Analysis," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2525-2536, Sept. 2016
  15. M. Rafferty, X. Liu, D. M. Laverty and S. McLoone, "Real-Time Multiple Event Detection and Classification Using Moving Window PCA," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2537-2548, Sept. 2016
  16. T. Jiang, Y. Mu, H. Jia, N. Lu, H. Yuan, J. Yan and W. Li, "A Novel Dominant Mode Estimation Method for Analyzing Inter-Area Oscillation in China Southern Power Grid," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2549-2560, Sept. 2016
  17. B. Wang, B. Fang, Y. Wang, H. Liu and Y. Liu, "Power System Transient Stability Assessment Based on Big Data and the Core Vector Machine," IEEE Transactions on Smart Grid, vol. 7, no. 5, pp. 2561-2570, Sept. 2016

Citation

Tao Hong, Chen Chen, Jianwei Huang, Ning Lu, Le Xie and Hamidreza Zareipour, "Guest Editorial: big data analytics for grid modernization", IEEE Transactions on Smart Grid, vol.7, no.5, pp 2395-2396, September, 2016

Guest Editorial: Big Data Analytics for Grid Modernization

Tao Hong, Chen Chen, Jianwei Huang, Ning Lu, Le Xie and Hamidreza Zareipour

Tuesday, July 26, 2016

Temperature Scenario Generation for Probabilistic Load Forecasting

When using weather scenarios to generate probabilistic load forecasts, a frequently asked question is
How many years of weather history do we need? 
This paper gives an answer based on an empirical study.

Most of my papers were accepted after two or more revisions. This time it only took one revision to have this paper accepted. In the first round of review, We received 40 comments from 6 reviewers. Our first revision was accepted after 4 of the reviewers recommended acceptance. In this blog post, I'm attaching the submitted version of the revision including our response letter. Some of our responses were rebuttals to one of the reviewers who made a personal attack on me.

Citation

Jingrui Xie and Tao Hong, "Temperature scenario generation for probabilistic load forecasting", Transactions on Smart Grid, accepted.

The working paper is available HERE.

Temperature Scenario Generation for Probabilistic Load Forecasting

Jingrui Xie and Tao Hong

Abstract

In today’s dynamic and competitive business environment, probabilistic load forecasting (PLF) is becoming increasingly important to utilities for quantifying the uncertainties in the future. Among the various approaches to generating probabilistic load forecasts, feeding simulated weather scenarios to a point load forecasting model is being commonly accepted by the industry for its simplicity and interpretability. There are three practical and widely used methods for temperature scenario generation, namely fixed-date, shifted-date, and bootstrap methods. Nevertheless, these methods have been used mainly on ad hoc basis without being formally compared or quantitatively evaluated. For instance, it has never been clear to the industry how many years of weather history is sufficient to adopt these methods. This is the first study to quantitatively evaluate these three temperature scenario generation methods based on the quantile score, a comprehensive error measure for probabilistic forecasts. Through a series of empirical studies on both linear and nonlinear models with three different levels of predictive power, we find that 1) the quantile score of each method shows diminishing improvement as the length of available temperature history increases; 2) while shifting dates can compensate short weather history, the quantile score improvement gained from the shifted-date method diminishes and eventually becomes negative as the number of shifted days increases; and 3) comparing with the fixed-date method, the bootstrap method offers the capability of generating more comprehensive scenarios but does not improve the quantile score. At the end, an empirical formula for selecting and applying the temperature scenario generation methods is proposed together with a practical guideline.