跳到正文
原文
Google AI:DEV 作者专属(RSS)· shubham goel·· 10 小时前AI 评分24

把 Platform API 变成真正的 Kubernetes 资源:Platform Operator 实现 Application 到 Deployment 与 Service 的转换

Turning My Platform API Into Real Kubernetes Resources

AI 导读

Platform Lab 新增 platform-operator 模块,将 Application 自定义资源(platform.shubforge.dev/v1alpha1)自动转换为 Deployment 和 Service,开发者只需定义 image、replicas 和 containerPort。

正文

In the previous post, I introduced the first real API for Platform Lab:

Application

The idea was to let an application developer define something simple like:

apiVersion: platform.shubforge.dev/v1alpha1
kind: Application

metadata:
  name: greeting-service

spec:
  image: greeting-service:1.0.0

  replicas: 1

  port:
    containerPort: 8080

At that point, Kubernetes could understand and validate the resource.

But nothing actually happened after creating it.

There was no Deployment.

There was no Service.

There was no platform behavior yet.

So the next step was to build the real Platform Operator.


What I Wanted to Build

The first platform flow is intentionally small.

Application
     |
     v
Platform Operator
     |
     +------ Deployment
     |
     +------ Service

The application developer should only define:

image
replicas
container port

The platform should handle the Kubernetes-specific resources.

For now, I deliberately avoided adding things like:

ConfigMaps
Secrets
health checks
resource requests
Istio
authorization
routes
observability

Those can come later when the platform actually needs them.


Creating the Platform Operator

Until now, I had a separate:

greeting-operator

which I used mainly to learn Kubernetes controller concepts.

For the actual platform work, I created a new module:

operators/platform-operator/

The idea is that this becomes the main operator for Platform Lab.

Eventually it can contain controllers for resources such as:

Application
ApplicationRelease
ApiDependency
AccessGrant

So I did not call it:

application-operator

because Application is only the first platform API.


Mapping the Application CRD to Java

The custom resource is represented in Java using:

Application.java
ApplicationSpec.java
ApplicationPort.java
ApplicationStatus.java

The main resource looks roughly like:

@Group("platform.shubforge.dev")
@Version("v1alpha1")
@Kind("Application")
@Plural("applications")
public class Application
        extends CustomResource<ApplicationSpec, ApplicationStatus>
        implements Namespaced {
}

This maps directly to:

apiVersion: platform.shubforge.dev/v1alpha1
kind: Application

The specification contains:

image
replicas
port

which matches:

spec:
  image: greeting-service:1.0.0

  replicas: 1

  port:
    containerPort: 8080

The First Real Translation

This is where the platform starts becoming useful.

The user writes:

spec:
  image: greeting-service:1.0.0
  replicas: 1

  port:
    containerPort: 8080

The Platform Operator translates that into:

Deployment
+
Service

So instead of thinking only in Kubernetes objects:

Deployment YAML
Service YAML
selectors
pod template
container ports
target ports

the developer works with a higher-level platform API.


Creating the Deployment

The Deployment is implemented as a managed dependent resource:

ApplicationDeploymentDependentResource.java

The operator takes the Application specification and builds a desired Deployment.

Conceptually:

Application
    |
    v
image
replicas
port
    |
    v
Deployment

The generated Deployment is roughly equivalent to:

apiVersion: apps/v1
kind: Deployment

metadata:
  name: greeting-service

spec:
  replicas: 1

  selector:
    matchLabels:
      platform.shubforge.dev/application: greeting-service

  template:

    metadata:
      labels:
        platform.shubforge.dev/application: greeting-service

    spec:

      containers:
        - name: application
          image: greeting-service:1.0.0

          ports:
            - containerPort: 8080

The application developer does not need to define any of this directly.


Creating the Service

The Service is another managed dependent resource:

ApplicationServiceDependentResource.java

The generated Service uses the same application label:

platform.shubforge.dev/application: greeting-service

The flow becomes:

Application
      |
      v
Deployment
      |
      v
Pods
      ^
      |
Service selector

The generated Service is roughly:

apiVersion: v1
kind: Service

metadata:
  name: greeting-service

spec:

  selector:
    platform.shubforge.dev/application: greeting-service

  ports:
    - name: http
      port: 8080
      targetPort: 8080
      protocol: TCP

This is the first place where the platform is hiding real Kubernetes wiring from the developer.


Why I Used Managed Dependent Resources

I could have written the controller manually like:

if Deployment does not exist
    create it

if Deployment exists
    update it

if Service does not exist
    create it

if Service exists
    update it

But that quickly becomes a lot of imperative logic.

Instead, I describe:

desired Deployment

and:

desired Service

and let the operator framework reconcile the actual Kubernetes resources toward that desired state.

So the mental model becomes:

Application spec
      |
      v
desired Deployment
      |
      v
actual Deployment

and:

Application spec
      |
      v
desired Service
      |
      v
actual Service

That feels much closer to the Kubernetes reconciliation model.


The Application Reconciler

The main reconciler is:

ApplicationReconciler.java

Its workflow includes both dependent resources:

Application
     |
     v
ApplicationReconciler
     |
     +----------------+
     |                |
     v                v
Deployment         Service

The dependent resources handle provisioning.

The main reconciler currently focuses mainly on status.


Initial Application Status

For the first version, I kept status simple.

The controller writes:

status:
  observedGeneration: 1
  deploymentName: greeting-service
  serviceName: greeting-service

This tells me:

which Application generation was processed

which Deployment belongs to it

which Service belongs to it

I intentionally did not set:

Ready=True

yet.


Why Ready Is Not Implemented Yet

Creating a Deployment does not mean the application is actually healthy.

For example, the Deployment may exist while Pods are failing with:

ImagePullBackOff
CrashLoopBackOff
failed readiness probe
scheduling failure

So:

resource created

is not the same thing as:

application ready

I want to implement readiness separately by actually observing Deployment status.

For this version, the operator only reports what it has provisioned.


Running the Platform Operator in Kubernetes

The operator itself is packaged as:

platform-operator:dev

and runs inside the local Kind cluster.

The Kubernetes deployment includes:

Namespace
ServiceAccount
ClusterRole
ClusterRoleBinding
Deployment

The operator runs as:

platform-system/platform-operator

RBAC

The first version only needs permissions for:

Applications
Application status
Deployments
Services

I did not copy every permission from the old Greeting operator.

The rule I want to follow is:

Add permissions only when the platform actually needs them.

So RBAC should grow with platform capabilities.


Creating the First Application

Once the operator was deployed, I created:

apiVersion: platform.shubforge.dev/v1alpha1
kind: Application

metadata:
  name: greeting-service

spec:
  image: greeting-service:1.0.0

  replicas: 1

  port:
    containerPort: 8080

Then checked:

kubectl get deployment greeting-service

and:

kubectl get service greeting-service

The platform had created both resources.

So the complete flow was finally working:

Application/greeting-service
            |
            v
      Platform Operator
        /          \
       v            v
Deployment       Service

This was the first point where Platform Lab started feeling like an actual platform rather than only an operator-learning project.


Testing the Controller

I also wanted testing to grow alongside the platform.

For the Application API itself, I already had:

application-api-test.sh

which checks CRD validation.

For the controller, I added:

application-controller-test.sh

This test uses the real local Kubernetes cluster and the real Platform Operator.

The flow is:

test script
     |
     v
temporary namespace
     |
     v
Application
     |
     v
Platform Operator
     |
     +------ Deployment
     |
     +------ Service

The test verifies that the platform really provisions the expected resources.


What the Controller Test Checks

The test creates an Application like:

spec:
  image: example/test:1.0.0

  replicas: 2

  port:
    containerPort: 8080

Then it waits for:

Deployment
Service

to appear.

After that it checks the Deployment.

For example:

replicas = 2
image = example/test:1.0.0
containerPort = 8080

Then it checks the Service:

port = 8080
targetPort = 8080

Finally, it verifies that Application status contains:

deploymentName
serviceName

So the test is proving the entire first platform flow.


Why I Started With a Shell Integration Test

I originally considered starting immediately with:

JUnit
Awaitility
Java Operator SDK test extensions

But for this stage, I wanted something easier to understand.

The current test is very explicit:

create Application
        |
        v
wait for Deployment
        |
        v
verify Deployment
        |
        v
wait for Service
        |
        v
verify Service

It talks to the same Kubernetes API I use manually.

That makes it easy to understand exactly what the test is proving.

Later I can still add more focused Java tests where they make sense.


Test Isolation

The controller test creates a temporary namespace.

Something like:

platform-test-12345

All test resources live there:

Application
Deployment
Service

At the end, the namespace is deleted.

So the test leaves the cluster clean.

The lifecycle is:

create temporary namespace
        |
        v
run test
        |
        v
delete namespace

Running the Tests

The Application API test is:

task application:test:api

The controller integration test is:

task application:test:controller

So testing is now becoming part of the normal platform development workflow.

My goal from this point is:

feature
   +
test
   +
documentation

rather than adding all tests later.


Current Platform Architecture

The current architecture now looks like:

                        Developer
                            |
                            v
                       Application
                            |
                            v
                    Platform Operator
                            |
              +-------------+-------------+
              |                           |
              v                           v
         Deployment                    Service
              |                           |
              v                           |
             Pods <-----------------------+

The Application resource is the source of desired state.

The Platform Operator converts that desired state into Kubernetes implementation details.


What This Version Does Not Handle Yet

There are still many things missing.

For example:

workload readiness
Application conditions
update behavior tests
ConfigMaps
Secrets
health checks
resource requests
routes
authorization
failure handling

I am intentionally leaving them out.

The current milestone is only:

Application
→ Deployment
→ Service

and proving that behavior works.


Source Code

The complete implementation is available in my Platform Lab repository.

Repository: Platform Lab

The changes covered in this post are available in:

Pull Request: Add Application Controller

The main additions include:

Platform Operator module
Application Java models
ApplicationReconciler
Deployment dependent resource
Service dependent resource
Platform Operator Kubernetes deployment
RBAC
Application controller integration test

What I Learned

The biggest shift in this step was moving from:

Kubernetes operator concepts

to:

platform abstraction

With the Greeting operator, my main question was:

How does a controller work?

With the Application controller, the question is becoming:

What Kubernetes complexity should the platform hide from an application developer?

The developer writes:

spec:
  image: ...
  replicas: ...
  port:
    containerPort: ...

and the platform generates:

Deployment
Service
selectors
Pod template
container port
service port

That translation is the real beginning of Platform Lab.


What's Next?

The next thing I want to test is whether the platform maintains desired state when the Application changes.

For example:

spec:
  image: greeting-service:2.0.0
  replicas: 3

should update the existing Deployment.

The flow should be:

Application generation 1
        |
        v
Deployment
image = 1.0.0
replicas = 1

        |
        | Application updated
        v

Application generation 2
        |
        v
same Deployment
image = 2.0.0
replicas = 3

That will let me explore Application updates and reconciliation as part of the real platform instead of only through the learning operator.

After that, I want to move toward:

Application readiness
configuration
Secrets
ApplicationRelease
ApiDependency
AccessGrant
networking
authorization

one capability at a time.

来源:Google AI:DEV 作者专属(RSS) · dev.to