跳到正文
arXiv:cs.LG· Yasaman Cheraghi (Department of Energy and Petroleum Engineering, University of Stavanger, Norway), Reidar B. Bratvold (Department of Energy and Petroleum Engineering, University of Stavanger, Norway), Aojie Hong (Independent Researcher, Stavanger, Norway), Ressi B. Muhammad (Department of Energy and Petroleum Engineering, University of Stavanger, Norway), Sergey Alyaev (NORCE Norwegian Research Centre, Bergen, Norway)·· 3 小时前

强化学习如何为能源转型中的价值创造做战略投资决策

Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach

AI 导读

挪威斯塔万格大学等机构的研究者构建了一个模拟至2050年动态能源格局的仿真环境,并设计多准则序贯决策框架,用于在油气、可再生能源和CO2减排三个板块间分配资金。研究评估用强化学习(RL)在该框架内寻找最优投资策略,智能体的序贯决策会影响油气产量、可再生能源出力、CO2排放和收入等关键变量。基准测试显示,RL策略在适应性和长期价值创造上持续优于人工设定的基线策略。

正文

Authors:Yasaman Cheraghi (1), Reidar B. Bratvold (1), Aojie Hong (2), Ressi B. Muhammad (1), Sergey Alyaev (3) ((1) Department of Energy and Petroleum Engineering, University of Stavanger, Norway, (2) Independent Researcher, Stavanger, Norway, (3) NORCE Norwegian Research Centre, Bergen, Norway)

View PDF

Abstract:The global challenge of climate change has driven significant steps to reduce CO2 emissions, guided by international agreements like the Paris Agreement of 2015. Acting too slowly could result in future losses and reputational damage, while moving too quickly could jeopardize shareholder value due to the marginal profitability or potential losses due to technology immaturity of many renewable projects. To navigate this complex transition, energy companies must adopt Sequential Decision Making (SDM) strategies to maximize value creation from decision flexibility under uncertainties. To support this, we developed a custom simulation environment to model the dynamic energy landscape up to 2050. Building on this, we designed a multi-criteria SDM framework that explores various decision strategies related to different portfolios for allocating funds across three sectors: oil & gas, renewables, and CO2 reduction. It aims to maximize value during the transition while accounting for uncertainties in productions, energy prices, and costs. This framework has three objectives: maximizing profit, minimizing CO2 social costs, and enhancing competitive advantage in the renewable energy sector. This research evaluates the use of Reinforcement Learning (RL) to identify optimal investment policies within the defined SDM framework. The agent's sequential decisions shape a virtual dynamic environment by influencing key variables such as oil and gas production, renewable energy output, CO2 emissions, and revenues. Through repeated interaction, the RL algorithm explores the state space and learns an optimal policy under uncertainty. We benchmark the RL strategy against a set of manually defined baseline policies and find it consistently outperforms them in adaptability and long-term value creation.
Subjects: Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG); Systems and Control (eess.SY)
Cite as: arXiv:2610.10768 [cs.CE]
  (or arXiv:2610.10768v1 [cs.CE] for this version)
  https://doi.org/10.48550/arXiv.2610.10768

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yasaman Cheraghi [view email]
[v1] Wed, 7 Oct 2026 18:29:31 UTC (1,903 KB)

来源:arXiv:cs.LG · arxiv.org