Friday, 18 November 2011

"Average Day" data plotting.

I have plotted an energy consumption of "average day" for different users. The results will be shown as follow:
Figure 1.0 - Average Day consumption for User ecenergy22 in 14 days
Figure 2.0 - Average Day consumption for User ecenergy23 in 11 days
Figure 3.0 - Average Day consumption for User ecenergy24 in 10 days
Figure 4.0 - Average Day consumption for User ecenergy25 in 7 days
Figure 5.0 - Average Day consumption for User ecenergy30 in 7 days
Figure 6.0 - Average Day consumption for User ecenergy32 in 18 days
Figure 7.0 - Average Day consumption for User ecenergy33 in 17 days
Figure 8.0 - Average Day consumption for User ecenergy34 in 11 days
Figure 9.0 - Average Day consumption for User ecenergy35 in 11 days
Figure 10.0 - Average Day consumption for User ecenergy36 in 12 days

Friday, 11 November 2011

Hourly consumption analysis

I have calculated an energy of user consumption in every hour to see if we can observe any patern, which could help to improve the prediction. I firstly run a whole data of user ecenergy22. The result can be seen in Figure 1.0 below:

Figure 1.0 User consumption in every hour.

Then, I generate the results in the first 13 days for an ease of observation. The following figures shows the daily user consumption in every hour unit.
Figure 2.1 User consumption on 06/09/2011.

Figure 2.2 User consumption on 07/09/2011.
Figure 2.3 User consumption on 08/09/2011.
Figure 2.4 User consumption on 09/09/2011.

Figure 2.5 User consumption on 10/09/2011. 
Figure 2.6 User consumption on 11/09/2011
Figure 2.7  User consumption on 12/09/2011.

Research Plan after prediction.

As promised, I rewrite my research plan here for your suggestion. So far, I have got GPs prediction running on the UK Carbon Intensity. The result of the UK carbon intensity can be acceptable. However, the prediction of the specific user consumption by GPs based on historical data is hard, probably impossible to formulate a reasonable covariance function. Hence, we need to use the annotated events from FigureEnergy system, (or even Non-Intrusive Load Monitoring (NILM) technique) to predict activities ahead, and improve the user demand prediction.

I assume those prediction, which mentioned above, could be done. After that, we focus on providing feedback such that users can be raised awareness of carbon intensity based on their everyday activities. Up to this point, we can have two options:

1 - From historical data, we can do some analysis and show information of devices usage and carbon intensity. From this information, we hope users can have more attention on their energy usage, so they can change their behaviour in a positive way.

2 - From the prediction, we can run optimisation to minimise the carbon intensity. Then, we can advise some action to users to reduce the carbon intensity in term of the use of their devices. Probably we can suggest users to defer some events from the high peak of the grid carbon intensity to other low peak.

Furthermore, we want to think of the feedback interface where users can colaborate with agents to plan their activities ahead.

Thursday, 10 November 2011

User consumption prediction analysis

Honestly it is really hard to predict the user consumption by using GPs. Previously, I have tried to run a few examples on the real user usage, unfortunately the result were not good. I think to be able to do that, we need some help from the users, who directly use their home devices.

People (or users) often have a plan of what they are going to do, typically day-ahead or week-ahead. Specifically, with the help of technology, they can do planning in their own calendar (e.g, google calendar), then synchonises all activities to their phone for the notification and better time management. Therefore, I think if we can access to this type of information, or if we can get users to support the prediction by working with the agent, we can predict more accurately. However, I am still not too sure how we can do that.

Back to the user consumption prediction analysis, I try to do some analysis on the labels, which were annotated by the users, to see if we can get something from there. I firstly imported the list of events from the excel file. In this file, it has the annotated labels for all users, then we need to filter the data of the specific user to do some test. After that, I sort the data in ascending order of the starting time of the event. Then, I do some calculation on labels so that each label has a starting time step t, running for a length of s. Each label contains an energy usage as well as a baseline usage, however we only want to focus on energy usage at this time.

Up to this stage, we have a list of labels in which each label has a name, a starting time step t, the length of a running time, and an usage. For example, kettle starts at time 3 to 7, with a consumption of 0.2023 kWh; washing machine runs from time 7 to 11, with a consumption of 0.305 kWh. I wonder how we can apply GPs to predict the future labels, as each label can has 3 parameters (label name, time, and consumption). One solution I can think of is to break down the list of labels to individual single category, for example under the label category such as kettle, washing machine,..., then apply GPs to predict time and usage. Hence, we aggregate all single prediction to get the final graph. I am not sure at this stage and looking for a suggestion.

In addition, Poison process can be used to calculate the probability of the labels, which should appear in the time step t'. But how can we combine the Poison process and GPs to make a good prediction, I still have not figure it out yet.

Wednesday, 2 November 2011

Total Daily Carbon Intensity in the UK - Plot and Prediction Analysis.

I try to analyse the data of Carbon Intensity in the UK. By getting the total daily carbon intensity in the UK, we can see the shape of the graph below (see figure 1).

Figure 1. Actual Total Daily Carbon Intensity in the UK (from 27 June 2011 to 27 September 2011)
As we can see in Figure 1, the carbon intensity during the weekday is apparently higher than the carbon intensity during the weekend.

After that, I apply a Single GPs prediction for the total daily carbon intensity in the UK. The initial training data set is the first 4 weeks (28 days), the predictive period starts from 29 to 92. In this prediction, I only apply one-day ahead prediction. For example, if we want to predict the carbon intensity for the day of 35, then I would consider the training data set is from 1 to 34. The result is in figure 2 below.

Figure 2. Prediction of Daily Total Carbon Intensity
The Mean Square Error (MSE) for this period of 64 days is 1944838. For more detail, Figure 3 shows the MSE for individual days during the predictive period.

Figure 3. MSE for Total Daily Carbon Intensity during the predictive period.

Let's look into further detail by predict Carbon Intensity for a day ahead, however we predict every half-an-hours instead of the total daily value. Previously, I only used the constant hyperparameter because it is time-consuming. This time, I use the initial training data set of 28 days, started from 27 June 2011. The predictive period is from the day of 29th to 58th. The hyperparameter is iterately trained everyday during the predictive period. Having waited a few hours (approximately 4 hours), Figue 4 below shows the result.

Figure 4. Single GPs on Carbon Intensity oneday ahead in the UK (with trained hyperparameters), started from 27 June 2011.
The MSE for this 30 days predictive period is 1000, which is very much improved. In addition, Figure 5.0 shows the MSE for this period in detail.

Figure 5.0. MSE of Single GPs on Carbon Intensity One Day Ahead, Every Half An Hour prediction.

Furthermore, we analyse an energy consumption of some real users. First of all, I plot the total daily of user's energy usage. These figures can be seen as following:

Figure 6.1. Daily Total Usage of User 1.
Figure 6.2. Daily Total Usage of User 2.
Figure 6.3. Daily Total Usage of User 3.
Figure 6.4. Daily Total Usage of User 4.
Figure 6.5. Daily Total Usage of User 5.

By using the same covariance function, which applied in Carbon Intensity prediction, I have tried to run some prediction. However, the result looks really bad. The covariance function for user's consumption has to be much different, which I still have not found out yet. Moreover, the resolution for the user's usage is every two minute, which is quite high. It typically takes much time to run the file and wait for the result. Particularly, when the training hyperparameter is applied, the waiting time could be take for a few hours.

I might need to reduce the resolution for the data of user's usage to do more test on GPs prediction.

Monday, 31 October 2011

Single GPs prediction for Carbon Intensity - Results

Taking the first 14 days as a training data set (from 27/06/2011 to 10/07/2011), I apply Single GPs to predict Carbon Intensity for the UK in the next 30 days (from 11/07/2011 to 10/08/2011). It generates a result, which can be seen from the picture below:
Single GPs for Carbon Intensity in the UK from 27/06/2011 to 11/08/2011.


I use the constant hyperparameter set of [-1.1112; -0.4744; 9.7377; -0.2624; -0.6790; -0.2624; -0.5132; -0.8945; -1.1858; -0.8945; -2.9447] to generate this result. The Mean Squared Error (MSE) for this period is 2233.

Moving on to the user consumption prediction, the file of user's consumption has two column: date&time, and the usage. The user usage is recorded every two minutes, as a result each day supposed to contain 720 values. However, there are some values missing in some specific days. For example, in 'ecenergy23@ecs.soton.ac.uk.csv', it only has 712 values on 09/09/2011 (8 values missing), and 719 values on 11/09/2011 (1 value missing). It could be more because I have not checked them all.

Thus, to apply GPs for user consumption, we firstly need to fill the missing value to make sure that it has a full 720 values for each day. Currently, the date & time column is in a format of 'dd/MM/yyyy hh:mm:ss'. We might need to convert the time, which is "hh:mm:ss" to the number of minutes for that day. The value now should be "dd/mm/yyyy A", where A = hh * 60 + mm (we do not consider seconds as it is always zero value). The x-axis now can be represented by A/1440.

Saturday, 29 October 2011

Single GPs prediction for Carbon Intensity.

Feeling a bit weird to come back and write this blog after having away for such a long time now, however it might be necessary to keep up my work. 

At the moment, I have spent time to understand and implement Gaussian Processes (GPs) to predict the UK's Carbon Intensity and User's consumption in particular. It is not easy to write a polished code in Matlab as I haven't used it before, so it also takes time to learn. 

Having struggled to apply a complicated covariance function in Matlab, I took a step back to read again the background of GPs from books and other's papers. I finally understand a bit more of how to correctly use GPs. As a result, I have done a Single-GPs predictions on  Carbon Intensity. Here is the results:

1 - My first attempt is to use GPs to predict Carbon Intensity in 30 days ahead.



2 - My second attempt is to use GPs to predict one day-ahead. Then, repeat the process until it predict up to 30 days. It looks better than the one-go 30 days prediction.

In these graphs, the solid blue line represents for the training data, which I took 14 days. The solid red line represents for the actual data, and the dashed black line represents for the predictive data.

To be updated...