As the next step in the codegen experiment, we want to generate the invoke calls for each layer. This is slightly challenging with the existing sources, as kernels only expose a registration function, not their individual Eval functions. In an effort to keep the code churn to a minimum, this PR introduces an inference only registration structure and function. It includes just two function pointers: invoke and reset. For this CL, we've only introduced it for FullyConnected.
In the code generator, this PR creates a new op_table array in the generated source, with an enum for lookup. It also generates an invoke function for each subgraph, that calls each operator's invoke function.
BUG=295174388
Change generated with the following command:
```bash
cd tensorflow/lite/micro/
sed -i 's/TfLiteRegistration/TfLiteRegistration_V1/g' `find . -name "*.h" -o -name "*.cc"`
cd -
cd third_party/hexagon/
sed -i 's/TfLiteRegistration/TfLiteRegistration_V1/g' `find . -name "*.h" -o -name "*.cc"`
cd -
tensorflow/lite/micro/tools/ci_build/test_code_style.sh --fix_formatting
```
BUG=http://b/272808609
* Fix linter issues in a few files.
* restore file whose source of truth is the upstream TF repo.
Co-authored-by: Advait Jain <advaitjain@google.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
* Enable -Wmissing-field-initializers and update all registration
references
In https://github.com/tensorflow/tflite-micro/pull/1082/, we disabled
the compiler error as a workaround for a successful upstream sync from
TFLite. This PR re-enables the error after addressing all errors. A new
`RegisterOp` API is added, and all prior TfLiteRegistration struct
initializations are updated to use the new API. At the minimum, this
ensures that such additions in future will only require an update to
RegisterOp if TFLM doesn't use that field.
BUG=b/230507399
* Added new RegisterOpWithFree for LSTM
* Remove LSTM Free
* Changed inline to normal function to save code size
* Changed inline to normal function to save code size
* Replace cmsis_nn static initialization with RegisterOp
* Removed extra code
* Add dropped namespace in hard_swish
Importing third_party/hexagon from codelinario at commit:
5bee22654219b89efdd35e447cc89d50369ad168
This is generated by running ci/import_third_party_hexagon.sh
BUG=https://b/227665919
* Replace GetInput/GetOutput with micro context api
* Replace GetInput/GetOutput in kernels with micro_context api. Ther are three kernels (strided_slice, squeeze, transpose)that are not changed yet because:
they put the returned value of GetInput/GetOutput in an OpData structure that is persistent.
It is better to have separate change to opdata of those kernels
* Fix corner case in pad
* Fix cmsis-nn
* Fix xtensa, cmsis_nn, hexagon
* Add Hexagon optimized kernels.
* Hexagon optimized kernels copied from https://source.codeaurora.org/quic/embedded_ai/tensorflow at 2d052806c211144875c89315a4fc6f1393064cf6
* Changed the include paths and directory structure a bit.
* Modified Makefile to allow optimized kernels to be in a separate directory
* Path to the Hexagon lib is now specified on the command line.
Verified that the optimized kernels are properly built and linked with:
```
make -f tensorflow/lite/micro/tools/make/Makefile TARGET=hexagon OPTIMIZED_KERNEL_DIR=hexagon OPTIMIZED_KERNEL_DIR_PREFIX=third_party HEXAGON_TFLM_LIB=~/Qualcomm/tflm_google/hexagon_tflm_core.a -j8 run_keyword_benchmark
```
Gives:
```
KeywordRunNIerations(1) took 52608 ticks (52 ms)
```
Whereas reference kernels with:
```
make -f tensorflow/lite/micro/tools/make/Makefile TARGET=hexagon -j8 run_keyword_benchmark
```
Gives:
```
KeywordRunNIerations(1) took 110248 ticks (110 ms)
```
BUG=http://b/190754463
* enable code style checks for third_party as well.
* clang-formatted the hexagon kernels.
* Rename to keep scope focused on what we currently need.