Saturday, 14 October 2017

Using Google Trend to Predict Stock Market Movement

Conventionally, financial and economic data such as Revenue, Profit, Debt, Return on Equity (ROE), Gross Domestic Product (GDP), Inflation etc., are analysed statistically to predict the stock market behaviour.  Lately, some traders are using Google Trend data to predict the stock market movement (Read more here).  Google Trend is a public web facility of Google Inc., based on Google Search, that shows how often a search-term is entered relative to the total search-volume across various regions of the world, and in various languages (Read more here).

One blogger noted that whenever celebrity Anne Hathaway’s name was mentioned in news, Warren Buffett’s Berkshire Hathaway shares rose (Read more here).  Is this purely coincidence?  Or some automated robot trading programs are taking place?

Let’s do a simple experiment, using Google Trend data to predict Kuala Lumpur Composite Index (KLCI).  Three search-terms – “Malaysia”, “1MDB”, and “KLCI” were selected.  The popularity of each search-term over time were plotted together with KLCI.  Chart 1 is search-term “Malaysia” and KLCI; Chart 2 is search-term “1MDB” and KLCI; while Chart 3 is search-term “KLCI” and KLCI.











Based on simple eye-balling inspection, Chart 1 and Chart 2 did not reveal any strong relationship between search-term and KLCI movement.  Although there was a sharp drop in KLCI when the popularity of “1MDB” surged in Aug 2015, the subsequent surge did not move KLCI drastically.  Chart 3, on the other hand, is more interesting as each time the popularity of search-term “KLCI” peaked, the KLCI tend to reverse its downtrend movement. 

Next, these data were analysed using basic machine learning algorithm.  Generally, there are two main types of machine learning used in quantitative finance – Regression, and Classification.  For simplicity purpose, Classification method is chosen for this analysis (Read more here).

The KLCI data were transformed into “Up”, “Down”, “Flat”, and “Dunno” by calculating the weekly closing price changes.  Example, if week 2 closing price is higher than week 1 closing price, week 2 will be classified as “Up”.  The “Down”, and “Flat” were calculated using similar said concept.  Additionally, the “Dunno” category was introduced to eliminate noises for the region where no high search popularity happened.

A time lag effect was also introduced into the model to “predict” whether KLCI will be “Up”, “Down”, “Flat”, or “Dunno” in the coming week.  As such, this week search-term results will affect next week KLCI behaviour.

Several algorithms were tested and k-nearest neighbours (KNN) algorithm was chosen as the accuracy is the highest among others.  See Picture 1 and Picture 2 for details.


Picture 1: Algorithm Comparison


 Picture 2: KNN Confusion Matrix and Accuracy Score


Now, let’s run a hypothetical test case to predict KLCI movement.  In hypothetical 1, assuming the search-term popularity for “Malaysia”, “1MDB”, and “KLCI” are 2, 1, and 25 respectively.  This means “Malaysia” and “1MDB” search traffics are almost flat but “KLCI” search traffic increase by 25%.  The KNN algorithm predicted the KLCI will go down in the following week.  In hypothetical 5, both “Malaysia and “1MDB” are almost flat but “KLCI” retreated from high peak.  The KNN algorithm predicated the KLCI will go up in the coming week.  The machine learning algorithm is giving similar results as eye-balling observation.  Table 1 shows KLCI movement predicted by KNN algorithm based on various hypothetical scenarios.

Table 1.

Above are just an illustrative example of how Google Trend and machine learning algorithm work.  Actual algorithm trading requires more intensive research and data processing effort! 

Saturday, 26 August 2017

Ringgit Forecast using Regression Analysis (Reblog)

According to Investopedia, there are six major factors that drive the value of a currency but not limited to - inflation, interest rate, current account balance, public debt, terms of trade, and political stability and economics performance (Read more here).

There could be other unique factors that influence the exchange rate.  For example, Malaysia is an oil and gas producer, one may wonder that whether oil price will have an impact on Ringgit exchange rate?  A simple method to study the relationship between Ringgit exchange rate and oil price, is the regression analysis.  Regression analysis is a statistical process for estimating the relationships among variables (Read more here).

The following graph shows the regression plot of USD/RM vs Brent Crude Oil from July 2005 to July 2017.  It shows that the Ringgit movement was significantly influenced by oil price.  Ringgit strengthened when oil price was high while the Ringgit weakened when oil price was low. 



The scattered dark blue dots are the actual data of USD/RM corresponding to respective Brent Crude Oil price from July 2005 to July 2015 while the scattered red dots are the actual data of USD/RM corresponding to respective Brent Crude Oil price from August 2015 to July 2017.  The light blue curve is the fitted regression line between USD/RM and Brent Crude Oil price (July 2005 – July 2017, R2 = 0.74) while the dashed red line is the + one standard deviation plot from the fitted regression line.

It is interesting to observe that before Aug 2015, the Ringgit was stronger and was well predicted by the regression line.  However, since Aug 2015, the Ringgit has weakened by + one standard deviation.

In the World Economic Outlook, April 2017 report, International Monetary Fund (IMF) predicted the oil price to be trading around USD55 per barrel in 2017 – 18 (Read more here).  Based on this, the Ringgit could be forecasted using the above regression analysis.  If all the post-2015 negative issues in Malaysia are resolved, Ringgit could trade around USD/RM 3.80.  Nevertheless, if the negative issues persist, the Ringgit may trade around USD/RM 4.18 as the above graph suggests.

Although regression analysis is widely used to infer causal relationships between the independent and dependent variables, do keep in mind that correlation does not imply causation!  The underlying causalities of the variables have to be analysed separately.


Reblog notice: The main content of this article was first published in MY MPCA blogspot (Read more here), where engineering2finance is the main author of that article.


Disclaimer:  The above analysis does not imply any buy or sell recommendation.  The author disclaims all liabilities arising from any use of the information contained in this article.




Friday, 4 August 2017

The Cost of Holding Physical Cash



The above photo shows a Malaysian 20 sen coin minted in 1967.  This 20 sen coin is still legit today and you can use it to buy something, perhaps, a piece of candy?

You may notice that you can’t really do much with this 20 sen nowadays because inflation has eroded the purchasing power.

Cash is king!  You always hear that.  But keeping cash without “transactions” is bad.  Cash shine because the “transactions” polish it!   Assuming you had invested this 20 sen in 1967 with a compounded annual return of 8%, it would worth RM 9.38 today, a whopping 46.9x of capital appreciation!  If you were not comfortable with the risk in investing, and opted for depositing this 20 sen in bank for 3% annual interested rate, it would worth 88 sen today.


But if you chose to keep this 20 sen in your drawer from 1967 to 2017, your penalty would be at least 3x loss!  In conclusion, keeping physical cash is a money losing practice!

However, if you keep the coin long enough, may be coin collector would like to offer you a good price for it!


Wednesday, 19 July 2017

Smart Beta Investing




I attended a luncheon and presentation on “Smart Beta” organized by CFA Society Malaysia on 22nd June 2017.  The speaker was Mr. Charles J. Yang, CFA, Managing Director of Tokio Marine Asset Management.  It was a great event and the speaker delivered a very concise presentation about Smart Beta Investing.

Mr. Yang pointed out that smart beta is not a “modified” mathematical definition of conventional beta.  It could be interpreted as a category of asset management, somewhere between passive and active management.  The total asset under management using smart beta approach was about US$ 416 billion in 2016.  Among the common factors of smart beta strategy such as “Low Volatility”, “High Dividend”, “Book Value”, “Cash Flow” and etc., Mr. Yang presentation was focusing on “Low Volatility”.

Generally, in strong macro environment, “Value”, “Momentum” and “Size” investing strategies will perform better while in weak macro environment, “High Dividend”, “Quality”, “Minimum Volatility” investing strategies will perform well.  Data shown that smart beta strategy has been performing well for the past 6 – 7 years.  In my opinion, the macro environment after the 2008 crash was neither strong nor weak, thus the combination of “Low Volatility”, “High Dividend” and “Value” of smart beta strategy is applicable.


As an engineering to finance apprentice, I asked some technical questions during the Q&A session.  My questions were:

  1. What was the time-period used to calculate beta?
  2. What was the data frequency used to calculate beta?
  3. Why these time-period and data frequency were chosen to calculate beta?
Below were Mr. Yang’s answers:

  1. Three years.
  2.  Monthly data.
  3. Trial and error basis using quantitative statistical method, need to back test.
From his reply we could infer that portfolio rebalancing is required periodically as the volatility of the assets in the portfolio may change from time to time, thus a quantitative team is essential to assist the smart beta strategy.