# `TFLiteElixir.LiteRT.CompiledModel.Server`
[🔗](https://github.com/cocoa-xu/tflite_elixir/blob/main/lib/tflite_elixir/litert/compiled_model_server.ex#L1)

A compiled model that lives inside a process, so several processes can share
one model without taking turns badly.

`TFLiteElixir.LiteRT.CompiledModel` refuses a second concurrent caller, which
is safe but leaves the caller to arrange the turns. This does the arranging:
calls are serialised by the process, and the model is claimed by it, so a
reference that escaped cannot be used behind its back.

    {:ok, env} = TFLiteElixir.LiteRT.CompiledModel.environment()
    {:ok, server} = TFLiteElixir.LiteRT.CompiledModel.Server.start_link(env, path)
    {:ok, outputs} = TFLiteElixir.LiteRT.CompiledModel.Server.run(server, inputs)

A model still runs one inference at a time, because LiteRT does; what the
process adds is that waiting is explicit and bounded rather than a race.

## The queue is bounded

A caller that submits faster than the model runs would otherwise grow the
mailbox until the node dies. Past `:max_queue` pending calls the server
answers `{:error, "the model's queue is full"}` instead, which is a back
pressure signal a caller can act on. The default is 64.

# `opts`

```elixir
@type opts() :: [{:max_queue, non_neg_integer()} | {atom(), term()}]
```

# `fully_accelerated`

```elixir
@spec fully_accelerated(pid()) :: {:ok, boolean()} | {:error, String.t()}
```

Whether the accelerator took the whole graph.

# `fully_accelerated!`

Raising version of `fully_accelerated/1`.

# `fully_accelerated?`

```elixir
@spec fully_accelerated?(pid()) :: boolean()
```

As `fully_accelerated/1`, answering `false` rather than an error.

# `io_sizes`

```elixir
@spec io_sizes(pid()) ::
  {:ok, {[non_neg_integer()], [non_neg_integer()]}} | {:error, String.t()}
```

The byte size of each input and output tensor, as `{inputs, outputs}`.

# `io_sizes!`

Raising version of `io_sizes/1`.

# `pending_events`

```elixir
@spec pending_events(pid()) :: {:ok, non_neg_integer()} | {:error, String.t()}
```

How many profiling events are waiting, without reading them.

# `pending_events!`

Raising version of `pending_events/1`.

# `profile`

```elixir
@spec profile(pid()) ::
  {:ok, [TFLiteElixir.LiteRT.CompiledModel.event()]} | {:error, String.t()}
```

Every profiling event recorded so far.

The profile belongs to the model, not to a call, so this covers every run any
process has made since the last reset.

# `profile`

```elixir
@spec profile(pid(), non_neg_integer()) ::
  {:ok, [TFLiteElixir.LiteRT.CompiledModel.event()]} | {:error, String.t()}
```

The most recent `limit` profiling events, or all of them when zero.

# `profile!`

Raising version of `profile/1`.

# `profile!`

Raising version of `profile/2`.

# `reset_profile`

```elixir
@spec reset_profile(pid()) :: :ok | {:error, String.t()}
```

Forget the events recorded so far and keep recording.

# `reset_profile!`

Raising version of `reset_profile/1`.

# `run`

```elixir
@spec run(pid(), [binary()]) :: {:ok, [binary()]} | {:error, String.t()}
```

Run the model over a list of input binaries.

# `run`

```elixir
@spec run(pid(), [binary()], timeout()) :: {:ok, [binary()]} | {:error, String.t()}
```

Run the model, waiting at most `timeout`.

# `run!`

Raising version of `run/2`.

# `run!`

Raising version of `run/3`.

# `run_with_metrics`

```elixir
@spec run_with_metrics(pid(), [binary()]) ::
  {:ok,
   {[binary()], [{binary(), TFLiteElixir.LiteRT.CompiledModel.metric_value()}]}}
  | {:error, String.t()}
```

Run and collect whatever counters the accelerator reports.

# `run_with_metrics`

```elixir
@spec run_with_metrics(pid(), [binary()], non_neg_integer()) ::
  {:ok,
   {[binary()], [{binary(), TFLiteElixir.LiteRT.CompiledModel.metric_value()}]}}
  | {:error, String.t()}
```

As `run_with_metrics/2`, at a given detail level.

# `run_with_metrics`

```elixir
@spec run_with_metrics(pid(), [binary()], non_neg_integer(), timeout()) ::
  {:ok,
   {[binary()], [{binary(), TFLiteElixir.LiteRT.CompiledModel.metric_value()}]}}
  | {:error, String.t()}
```

As `run_with_metrics/3`, waiting at most `timeout`.

# `run_with_metrics!`

Raising version of `run_with_metrics/2`.

# `run_with_metrics!`

Raising version of `run_with_metrics/3`.

# `run_with_metrics!`

Raising version of `run_with_metrics/4`.

# `start`

```elixir
@spec start(reference(), String.t()) :: {:ok, pid()} | :ignore | {:error, term()}
```

Start one outside a supervision tree.

# `start`

```elixir
@spec start(reference(), String.t(), opts()) ::
  {:ok, pid()} | :ignore | {:error, term()}
```

Start one outside a supervision tree, with options.

# `start_link`

```elixir
@spec start_link(reference(), String.t()) :: {:ok, pid()} | :ignore | {:error, term()}
```

Start a compiled model process linked to the caller.

# `start_link`

```elixir
@spec start_link(reference(), String.t(), opts()) ::
  {:ok, pid()} | :ignore | {:error, term()}
```

Start a compiled model process linked to the caller.

Takes everything `TFLiteElixir.LiteRT.CompiledModel.new/3` takes, plus:

- `:max_queue`. How many calls may be waiting before further ones are refused.
  Defaults to 64.

# `stop`

```elixir
@spec stop(pid()) :: :ok
```

Stop the process, and with it the compiled model.

# `summarise_profile`

```elixir
@spec summarise_profile(pid()) ::
  {:ok, [TFLiteElixir.LiteRT.CompiledModel.summary_entry()]}
  | {:error, String.t()}
```

Per-operator totals over every run since the last reset, slowest first.

# `summarise_profile!`

Raising version of `summarise_profile/1`.

# `with`

```elixir
@spec with(pid(), (reference() -&gt; result)) :: result | {:error, String.t()}
when result: term()
```

Run a function against the compiled model inside the owning process.

The escape hatch for anything this module does not forward. The reference is
only usable for the duration of the call, because the server owns the model
and takes it back afterwards. A function that raises costs the call and not
the model.

# `with`

```elixir
@spec with(pid(), (reference() -&gt; result), timeout()) :: result | {:error, String.t()}
when result: term()
```

As `with/2`, waiting at most `timeout`.

---

*Consult [api-reference.md](api-reference.md) for complete listing*
