Predicting Stock with Deep RL

Updated on Dec 26,2023

Predicting Stock with Deep RL

Table of Contents:

  1. Introduction
  2. Importing Libraries
  3. Loading and Preprocessing Data
  4. Creating the Stock Trading Environment
  5. Training the Deep Q-Network (DQN)
  6. Evaluating the Trained Model
  7. Visualizing the Training Process
  8. Conclusion

Introduction

In this article, we will explore how to predict stocks using deep reinforcement learning. We will be coding in Python using Jupyter Notebook, and we will use a dataset called the "Huge Stock Market Dataset" for our predictions. If You want to follow along and understand the complete explanation, you can download the source code from the link provided in the description. Let's dive in and start coding!

Importing Libraries

The first step in our project is to import all the necessary libraries. We will be using libraries such as time, copy, numpy, pandas, pytorch, shanen, and plotly. Each library serves a specific purpose in our deep reinforcement learning project. Let's go through each import and its purpose:

  1. Importing the time library: This library provides various time-related functions which might be useful for measuring the time taken by our code or for other timing purposes.

  2. Importing the copy library: The copy module is used for creating shallow and deep copies of objects. This is useful when we want to duplicate Mutable objects without changing the original.

  3. Importing the numpy library: Numpy is a widely used library for numerical computing in Python. It provides efficient and convenient manipulation of arrays and other numerical data structures. We import it with the alias np for easier access.

  4. Importing the pandas library: Pandas is a popular library for data manipulation and analysis. It provides data structures like data frames and series, which are useful for handling tabular data. We import it with the alias pd for easier access.

  5. Importing pytorch: PyTorch is a deep learning framework that allows you to write flexible and efficient neural network models. It is designed to be easy to use and provides automatic differentiation capabilities.

  6. Importing shannon.functions as F: This imports the function submodule of shannon, which provides various activation functions, loss functions, and other operations for deep learning. We import it with the alias F for easier access.

  7. Importing shannon.links as L: This imports the Links submodule of shannon, which provides predefined neural network layers such as fully connected layers, convolutional layers, and recurrent layers. We import it with the alias L for easier access.

  8. Importing plotly.tools: This imports the tools submodule of the plotly library, which is a graphing library for making interactive and high-quality graphs. The tools submodule contains various utility functions for working with plotly charts.

  9. Importing plotly.graph_objects: This imports all the classes from the graph_objects submodule of plotly, which contains objects for creating various types of plots such as scatter plots, line plots, and bar plots.

  10. Importing plotly.offline as plt: This imports the init_notebook_mode function from the offline submodule of plotly, which allows us to Create and display interactive plots within a Jupyter Notebook without needing an internet connection.

  11. init_notebook_mode(): This function initializes the notebook mode for plotly, which means that the interactive plots will be rendered within the Jupyter Notebook itself.

By importing these libraries and initializing the notebook mode, we have set up our environment to work with stock data, build deep reinforcement learning models using shanen, and Visualize the results with plotly.

Loading and Preprocessing Data

The next step in our project is to load the stock data and preprocess it using the pandas library. We will split the data into training and testing sets, convert the date column to the Pandas datetime format, and set the date column as the index of the data frame for easier manipulation and analysis. Let's go through each step of this process:

  1. Loading the data

    • We use the pd.read_csv() function to Read the CSV file containing the stock data for Google and store it in a pandas data frame called data. The file is located in the relative path ./input/Data/Stock/google.us.txt.
  2. Preprocessing the data

    • Converting the date column: We use the pd.to_datetime() function to convert the date column of the data frame from a STRING format to Pandas datetime format. This allows for easier manipulation of date-related operations.
    • Setting the date as the index: We use the data.set_index() function to set the date column as the index of the data frame. Using the date as the index allows us to easily filter and perform time series analysis on the data.
    • Printing the minimum and maximum dates: We use the data.index.min() and data.index.max() functions to print the minimum and maximum dates in the data set. This gives us an idea of the time range covered by the data.
    • Printing the first five rows: We use the data.head() function to print the first five rows of the data frame. This is useful for getting a quick look at the data and making sure it has been loaded and processed correctly.

By following these steps, we have loaded the Google stock data into the data data frame, converted the date column to the Pandas datetime format, and set the date as the index of the data frame. We also printed the minimum and maximum dates to get an idea of the time range covered by the data, and we printed the first five rows to verify that the data has been loaded and processed correctly.

Creating the Stock Trading Environment

To train our deep reinforcement learning model, we need to create a stock trading environment that simulates the process of trading stocks. This environment will allow our agent to Interact with the stock data, take actions (buy, sell, or hold), and receive rewards Based on the profitability of its actions. The environment will be implemented as a class called Environment, which will have methods for resetting the environment, taking actions, and returning the Current state, reward, and done flag. Let's go through each step of creating the stock trading environment:

  1. Class Definition: We define a class called Environment, which will be a subclass of the gym.Env class from the OpenAI Gym library. This allows us to leverage the Gym environment API for creating custom environments.

  2. Initialization: The constructor of the Environment class takes two arguments: data and history_T. data is the pandas data frame containing the stock data, and history_T is the number of recent time steps to consider for the state representation. We store these arguments as instance variables.

  3. Reset Method: The reset() method resets the environment to its initial state. It sets the time step to zero, sets the done flag to False (indicating that the environment is not finished), resets the total profits to zero, and resets the list of open positions.

  4. Step Method: The step() method takes an action (act) as input and performs one step in the environment. It updates the state, calculates the reward, and returns the new observation, reward, and done flag. The possible actions are:

    • 0: Hold (do nothing)
    • 1: Buy (open a new position at the current close price)
    • 2: Sell (close all positions and calculate the profit)
  5. Profit Calculation: To calculate the profit, we iterate over all the open positions and calculate the profit for each position based on the difference between the close price and the price at which the position was opened.

  6. Position Update: The update_positions() method updates the current value of all open positions by multiplying the position quantity by the close price of the current time step.

  7. Observation Calculation: The observation at each time step is represented as a list containing the position value and a list of price differences. The length of the list of price differences is equal to history_T.

  8. Reward Clipping: To avoid large rewards or losses, we clip the reward to -1, 0, or 1 depending on whether the action resulted in a loss, no change, or a gain, respectively.

By implementing these steps, we have created a stock trading environment that simulates the process of trading stocks. The environment allows our agent to take actions (buy, sell, or hold), receive rewards based on the profitability of its actions, and observe the state of the market.

Training the Deep Q-Network (DQN)

Once we have our stock trading environment, we can train a Deep Q-Network (DQN) on the training data to learn a trading strategy. The DQN will take the current observation as input, predict the best action to take, and update its Q-values based on the observed rewards. The training process consists of iterating over multiple episodes, where each episode represents a complete pass through the training data. Let's go through each step of training the DQN:

  1. DQN Network Architecture: We define the architecture of the DQN by creating a class called DQN_Network that represents the neural network model. The model consists of three fully connected layers (fc1, fc2, and fc3) with ReLU activation functions for the Hidden layers.

  2. Hyperparameters: We set various hyperparameters for the training process, such as the number of training epochs, the maximum number of steps per epoch, the size of the replay memory, the batch size for training, the exploration rate (epsilon), and the discount factor for future rewards (gamma).

  3. Training Loop: We iterate over the training epochs. For each epoch, we reset the environment, select an action based on the epsilon-greedy exploration strategy, take the selected action, update the Q-values, and calculate the reward and loss for each step. After a certain number of steps, we update the target network and adjust the epsilon value.

  4. Target Network Update: We update the target Q-Network periodically by copying the weights from the main DQN. This helps stabilize the training process and improves the convergence of the Q-values.

  5. Loss Calculation: We calculate the loss using the mean squared error (MSE) loss function. The loss measures the discrepancy between the predicted Q-values and the target Q-values.

  6. Model Evaluation: After training the DQN, we can evaluate its performance on the testing data set. We use the trained model to simulate trading decisions on the testing data set, calculate the profits, and compare them to a baseline strategy.

By following these steps, We Are able to train a DQN on the training data set and evaluate its performance on the testing data set. The trained DQN can be used to make trading decisions based on the learned strategy and potentially improve the profitability of the baseline strategy.

Evaluating the Trained Model

Once the DQN model is trained, we can evaluate its performance on the testing data set by simulating trading decisions using the learned strategy. We compare the profits generated by the model with the profits generated by a baseline strategy to assess the effectiveness of the learned trading strategy. Let's go through the steps of evaluating the trained model:

  1. Simulating Trading Decisions: We loop through all the data points in the testing set and use the trained DQN to predict the action to take at each time step. We add the predicted action to the list of actions taken.

  2. Calculating Profits: We calculate the profits generated by the DQN model by getting the actions taken, positions, and prices from the environment.

  3. Comparing with Baseline Strategy: We compare the profits generated by the DQN model with the profits generated by a baseline strategy, such as a buy-and-hold strategy or a random trading strategy.

  4. Visualizing the Trading Actions: We create a plot that visualizes the trading actions taken by the DQN agent on both the training and testing data sets. The plot shows the stock price movement along with the corresponding actions, where different colors represent different actions (gray for hold, cyan for buy, and magenta for sell).

By following these steps, we can assess the performance of the trained DQN model and compare it with a baseline strategy. Visualizing the trading actions provides insights into how the model makes decisions and how well it adapts to different market conditions.

Visualizing the Training Process

To understand how the model learns and improves over time, we can visualize the training process by plotting the training losses and rewards. This allows us to track the convergence of the model's Q-values and assess its learning progress. Let's go through the steps of visualizing the training process:

  1. Creating the Plot: We create a figure with two subplots, one for the losses and one for the rewards. Each subplot contains a line Chart with the x-axis representing the training epochs and the y-axis representing the loss or reward values.

  2. Adding the Loss and Reward Traces: We add the loss and reward traces to their respective subplots using the scatter function from the plotly.graph_objects module. The line colors are set to Sky Blue for the loss and orange for the reward.

  3. Setting the Plot Titles: We set the x-axis titles of the subplots to "Epoch" and update the layout of the plot.

  4. Displaying the Plot: We display the plot using the plt.Show() function.

By visualizing the training losses and rewards, we can gain insights into the learning process of the model. Ideally, we would like to see the loss values decreasing over time, indicating that the model is learning to approximate the optimal Q-values. Similarly, observing an increasing trend in the reward values would suggest that the agent is learning a profitable trading strategy.

Conclusion

In this article, we have explored the process of predicting stocks using deep reinforcement learning. We have covered various steps, including importing libraries, loading and preprocessing data, creating the stock trading environment, training the DQN, evaluating the trained model, visualizing the training process, and summarizing the main findings. By following these steps, you can implement your own deep reinforcement learning model for stock prediction and explore the possibilities of algorithmic trading.

Most people like